Qwen 3.6 35B A3B
Alibaba's open-weight MoE coding specialist — 35B total, 3B active, built for agentic development and long-context reasoning.
Get API keyWhat is Qwen 3.6 35B A3B?
Qwen 3.6 35B A3B is an open-weight Mixture-of-Experts text model from Alibaba's Qwen team, released in April 2026. It routes 35 billion total parameters through 3 billion active per token, prioritizing agentic coding, reasoning, and long-context workflows under an Apache 2.0 license.
Use Qwen 3.6 35B A3B privately on Venice
On Venice, Qwen 3.6 35B A3B runs under a private, zero-retention tier — your prompts are not stored, profiled, or used for training. You get access to its tool use, reasoning, web search, and structured output capabilities with no subscription, paying only $0.15 per 1M input tokens and $1 per 1M output tokens.
What can Qwen 3.6 35B A3B do?
- •Open-weight Apache 2.0 model with MoE efficiency — only 3B parameters are activated per token, lowering inference cost while maintaining 35B-scale capacity.
- •Built for agentic coding and complex reasoning, with native support for tool use, web search, structured JSON output, and reasoning modes on Venice.
- •Supports reasoning modes that preserve thinking context across conversation turns, streamlining iterative development and debugging workflows.
- •256K-token context window and up to 65,536 output tokens on Venice, with FP8 quantization for fast inference.
- •Extremely cost-efficient token pricing on Venice ($0.15/1M input) compared to closed frontier models.
- •Not uncensored — the model includes standard safety alignment and may refuse certain sensitive or restricted requests.
- •While the open weights include a vision encoder, Venice currently exposes text-only generation for this model.
- •As an MoE model, it requires compatible inference infrastructure; however, Venice handles this transparently.
- •Benchmark comparisons with dense or closed rivals are inherently imprecise because evaluation protocols (scaffolding, temperature, thinking mode) differ widely.
- •Output token cost ($1/1M) is higher than some rival open-weight models, though input pricing is very low.
Qwen 3.6 35B A3B capabilities
- Tool use / function calling
- Vision (image input)
- Reasoning
- Web search
- Code-optimized
- Structured output (JSON schema)
- Audio input
- Video input
- Multiple image inputs
- Log probabilities
How to use Qwen 3.6 35B A3B via API
Venice exposes an OpenAI-compatible API. Swap your base URL and call qwen3-6-35b-a3b.
curl https://api.venice.ai/api/v1/chat/completions \
-H "Authorization: Bearer $VENICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-6-35b-a3b",
"messages": [{ "role": "user", "content": "Explain quantum tunneling simply." }]
}'Specifications
Pricing
Billed per token on Venice: $0.15 per 1M input tokens and $1 per 1M output tokens.
New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.
Qwen 3.6 35B A3B vs alternatives
| Model | Context window | Open weights | Price (Venice) | Best for |
|---|---|---|---|---|
| Qwen 3.6 35B A3B | 256K tokens | Yes | $0.15 in · $1 out / 1M | Agentic coding & reasoning |
| DeepSeek V3.2 | 160K tokens | Yes | $0.33 in · $0.48 out / 1M | General MoE efficiency |
| Google Gemma 4 31B Instruct | 256K tokens | Yes | $0.12 in · $0.36 out / 1M | Dense open-weight tasks |
| Kimi K2.6 | 256K tokens | Yes | $0.75 in · $3.50 out / 1M | Long-context reasoning |
An efficient open-weight MoE with thinking preservation and strong coding performance.
What is Qwen 3.6 35B A3B good for?
- •Agentic software engineering — repo-level reasoning, frontend workflows, and multi-file coding.
- •Backend API and tool-augmented development with function calling and structured JSON output.
- •Long-document analysis, summarization, and extraction across 256K-token contexts.
- •Iterative debugging assistants that leverage reasoning and context preservation across turns.
- •Research and automation pipelines that integrate web search.
Prompting tips
- •Use explicit reasoning instructions like 'Think step by step' to engage the model's reasoning mode.
- •Leverage function calling and JSON schema for agentic pipelines and structured data extraction.
- •Place key instructions at the top or bottom of long prompts to maximize recall across the 256K context window.
- •Iterate across turns; the model preserves reasoning context well, reducing overhead on follow-up requests.
Version history
Predecessor series that preceded the Qwen3.6 architecture.
Dense variant with 27B active parameters.
CurrentCurrent MoE variant — 35B total, 3B active.
Frequently asked questions
Qwen 3.6 35B A3B is an open-weight Mixture-of-Experts text model from Alibaba's Qwen team. It has 35 billion total parameters with 3 billion active per token, built for agentic coding, reasoning, and long-context tasks under the Apache 2.0 license.
Venice charges $0.15 per 1M input tokens and $1 per 1M output tokens. Cached input costs $0.05 per 1M tokens. There is no subscription; you pay per token used.
Yes. The weights are released under the Apache 2.0 license, permitting commercial use, modification, and self-hosting. On Venice, the model runs privately with zero retention of your prompts.
Yes. On Venice, the model supports function calling, structured JSON output, reasoning, and web search, making it suitable for agentic coding and automation pipelines.
Choose Qwen 3.6 35B A3B for agentic coding, thinking preservation, and a 256K context window. Choose DeepSeek V3.2 if you prefer lower output pricing and general MoE efficiency, keeping in mind its 160K context limit.
No. It is not uncensored and includes standard safety alignment. It may decline certain sensitive, harmful, or restricted requests.
On Venice, the model supports a 256K-token context window and up to 65,536 output tokens per generation.
Yes. Venice runs it under a private tier with zero retention — your prompts are not stored, profiled, or used for model training.
Related models
Run Qwen 3.6 35B A3B privately.
No prompt logging. No data used for training. Free to start — no credit card.
