Now on VeniceLLMReasoningAnonymous

Abliterated Large V2

GLM 5.3 with refusal directions removed from the weights — built for offensive cyber, red-teaming, and long-horizon agent testing that other APIs shut down mid-run.

For agents
curl https://api.venice.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "abliteration-abliterated-model-large-v2",
    "messages": [{ "role": "user", "content": "Explain quantum tunneling simply." }]
  }'
Model IDabliteration-abliterated-model-large-v2
Maker
Abliteration AI (US)
Context
1,000K tokens
Reasoning
Supported
Privacy
Anonymous

Overview

What is Abliterated Large V2

Abliterated Large V2 is Abliteration AI's refusal-removed version of Z.ai's GLM 5.3, released August 29, 2026. Abliteration strips the refusal directions from the weights rather than jailbreaking prompts, so the model completes authorized vulnerability research, exploit development, and red-team agent work that the base model declines — with coding and agentic performance intact.

Using it anonymously on Venice

On Venice, Abliterated Large V2 runs under the anonymized privacy tier: your prompts are not stored, profiled, or used for training — a meaningful pairing for security work where engagement scope and client data are sensitive. You get per-token billing instead of a subscription, plus tool use, reasoning, web search, and structured JSON output through the OpenAI-compatible Venice API. Note that Venice does not run this model in a TEE or with end-to-end encryption; privacy comes from zero retention of your prompts.

AnonymousNo prompt trainingTEE · hardware enclaveEnd-to-end encrypted

Agent quickstart

Three calls, copied straight out

The API is OpenAI-compatible: change the base URL and the model id and existing client code works unchanged.

Streaming chat

curl https://api.venice.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "abliteration-abliterated-model-large-v2",
    "stream": true,
    "messages": [{ "role": "user", "content": "Draft the release note." }]
  }'

Tool calling

curl https://api.venice.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "abliteration-abliterated-model-large-v2",
    "messages": [{ "role": "user", "content": "Find the rate limits." }],
    "tools": [{
      "type": "function",
      "function": {
        "name": "search_docs",
        "description": "Search the API documentation.",
        "parameters": {
          "type": "object",
          "properties": { "query": { "type": "string" } },
          "required": ["query"]
        }
      }
    }],
    "tool_choice": "auto"
  }'

Python SDK

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["VENICE_API_KEY"],
    base_url="https://api.venice.ai/api/v1",
)

resp = client.chat.completions.create(
    model="abliteration-abliterated-model-large-v2",
    messages=[{"role": "user", "content": "Build without permission."}],
)
print(resp.choices[0].message.content)

Specifications

Datasheet

Maker
Abliteration AI (US)
Modality
Text only — no image, audio, video, or file input
Open weights
No — the abliterated build is proprietary to Abliteration AI; the underlying GLM 5.3 base is open-weight
License
Proprietary (abliterated build of open-weight GLM 5.3)
Modes
reasoning_effort parameter from none to max (thinking budget)
Context window
1,000K tokens
Input images
Not supported
Released
August 29, 2026
Base model
Z.ai GLM 5.3, abliterated post-training (refusal directions removed from the weights)
Hosting
FP8, US-hosted
Max output
32K tokens
Capabilities
Function calling, Reasoning, Web search
Privacy on Venice
Anonymized — prompts not stored
Available on Venice since
Oct 2026

Assessment

Strengths and limitations

Strengths
  • Purpose-built for the jobs other APIs refuse: authorized offensive cyber, AI red-teaming, trust-and-safety adversarial testing, and exploit-chain work that stalls on refusal layers elsewhere.
  • Strong agentic and coding benchmarks for its class — maker-reported 84.5% pass@1 on CyberGym (1,507 OSS-Fuzz bugs) and 41.8% on Terminal-Bench 4.0.
  • Full agent stack on Venice: tool use / function calling, reasoning with an adjustable effort budget, web search, and strict structured output (JSON schema).
  • 1M-token context window handles entire codebases, long engagement logs, and multi-hour agent runs in a single session.
  • Prompt caching bills cached input at 10% of the standard rate ($0.30 per 1M), which is cheap for repeated system prompts and tool definitions.
  • Drop-in OpenAI/Anthropic-compatible endpoint: swap the model name, keep your harness.
Limitations
  • Closed build: the abliterated weights are not published, so you cannot self-host or fine-tune them; only the GLM 5.3 base is open.
  • Not the benchmark leader across the board: maker-reported numbers put GPT-5.5 ahead on CyberGym (85.6%) and Claude Opus 5 ahead on Terminal-Bench 4.0 (51.8%).
  • Text-only: no image, audio, or document input, which limits multimodal red-team workflows.
  • Abliteration is a permanent weight modification, not a policy layer: governance, scoping, and authorization sit entirely on your side of the API call.
  • The approach is contested: security journalists have demonstrated the same access can produce clearly malicious output, so expect scrutiny if you deploy it in regulated environments.

Use cases

What it is good for

  1. 01Authorized penetration-testing agents that need to finish an exploit chain without a refusal layer killing the run.
  2. 02AI red-teaming: evaluating other models' safety behavior with an attacker model that doesn't self-censor.
  3. 03Vulnerability reproduction against OSS-Fuzz-style bug corpora at scale.
  4. 04Long-horizon coding agents working across entire repositories in one context window.
  5. 05Trust-and-safety adversarial prompt testing for platforms and labs.

Prompting

Getting better results

Set reasoning_effort explicitly: low for quick triage, max for multi-step exploit chains — the thinking budget is a real dial, not a fixed behavior.

Pass a JSON schema for structured output so findings, PoCs, or triage verdicts come back machine-parseable instead of prose.

Define your tools (shell, scanner, sandbox) via function calling and let the model drive them — this model is tuned for agentic loops, not one-shot Q&A.

Front-load your engagement scope and rules of engagement in the system prompt, then keep that prefix stable so prompt caching bills it at $0.30 per 1M tokens.

Feed whole repos or long logs into the 1M-token context rather than summarizing — the model resolves long-horizon tasks better with full source.

State the authorization context (client, scope, dates) in the prompt; the refusal layer is gone, but precise scoping still sharpens output quality.

Alternatives

How it compares

ModelBest forContextOpen weightsPrice (Venice)
Abliterated Large V2Offensive cyber, red-teaming & agent testing1M tokensNo$3 in · $5 out / 1M
Claude Opus 5Frontier reasoning & general coding1M tokensNo$6 in · $30 out / 1M
DeepSeek V4.1 FlashHigh-volume coding on a budget1M tokensYes$0.38 in · $1.50 out / 1M
Kimi K3Open-weight long-context agents1M tokensYes$3.75 in · $18.75 out / 1M

Abliterated Large V2 is the right pick for authorized offensive-security, red-team, and adversarial-testing agents — the niche where refusal layers, not capability, are the bottleneck. For general frontier coding without that constraint, Claude Opus 5 or DeepSeek V4.1 Flash are stronger value plays.

Pricing

What it costs on Venice

Billed per token on Venice: $3 per 1M input tokens and $5 per 1M output tokens.

Input / 1M tokens
$3
Per 1M tokens
Output / 1M tokens
$5
Per 1M tokens
Cached input / 1M
$0.30
Per 1M tokens

New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.

Getting a key

From nothing to a first call

  1. 01

    Create a key in API settings. Nothing else is required to start.

  2. 02

    Export it as VENICE_API_KEY so the snippets above run unedited.

  3. 03

    Point an existing OpenAI client at https://api.venice.ai/api/v1. The scheme is part of the value: an OpenAI client given a bare host does not resolve it.

  4. 04

    Pass abliteration-abliterated-model-large-v2 as the model and send the request.

FAQ

Frequently asked questions

Abliterated Large V2 is Abliteration AI's refusal-removed build of Z.ai's open-weight GLM 5.3, released August 29, 2026. Abliteration identifies the refusal direction in the model's hidden states and removes it at the weight level — a permanent modification, not a jailbreak — so the model completes authorized offensive-cyber, red-team, and agent-testing tasks the base model declines, while keeping GLM 5.3's coding and agentic performance.

Venice bills Abliterated Large V2 per token: $3 per 1M input tokens and $5 per 1M output tokens, with cached input at $0.30 per 1M. There is no subscription required — you pay per token from your credit balance.

Neither, strictly. Using it on Venice requires credits, though new accounts include free daily prompts and welcome credits to try it. On weights: the abliterated build itself is proprietary to Abliteration AI and not published, but it is derived from GLM 5.3, which is open-weight — so the base is open, the refusal-stripped version is not.

The model was built by removing refusal behavior from the weights, so it follows through on offensive-security and red-team tasks that standard models decline. Venice lists it as a standard text model rather than tagging it uncensored, and your own usage policy — not a provider's — sits on top. You are responsible for authorization and scope.

Abliteration AI reports 84.5% pass@1 on CyberGym (1,507 OSS-Fuzz bugs across 188 projects), 41.8% on Terminal-Bench 4.0, and 105 of 869 ExploitGym tasks completed in a two-hour window. These are maker-reported numbers; competitors score higher on some suites, so treat them as a profile of the model's target workload rather than a clean sweep.

Yes. On Venice it supports function calling, an adjustable reasoning effort (none to max), web search, and strict structured output via JSON schema — the full stack needed to drive penetration-testing or red-team agent harnesses through the OpenAI-compatible API.

Claude Opus 5 is the stronger general frontier model and costs more ($6 in · $30 out per 1M vs $3/$5). Abliterated Large V2 exists for the opposite case: when Opus-class models find the vulnerability but refuse to finish the exploit chain or red-team eval. If your work is standard coding and reasoning, choose Opus 5; if refusals are killing your runs, choose Abliterated Large V2.

Venice runs it under the anonymized privacy tier: prompts are not stored, tied to a personal profile, or used for training. Venice does not execute this model inside a trusted execution environment or with end-to-end encryption, so the guarantee is zero retention of your prompts rather than hardware-level isolation.

Z.ai's GLM 5.3, an open-weight frontier model. Abliteration AI applies its abliteration post-training to strip refusal directions from the weights and hosts the result in FP8 on US infrastructure. The base model's coding and agentic capabilities carry over; only the refusal behavior changes.

Use Abliterated Large V2 anonymously

One key, free to start, no credit card.

Start chat