← All models
Online

Qwen3.8-2.4T-A95B

Qwen · OpenAI-compatible API

qwen/qwen3.8-2.4t-a95b

USD / 1M tokens

Input$1.95
Cached Input$0.245
Output$5.70
Context
1M
Tool Calling
✓
Reasoning
✓
Vision
✓
Streaming
✓
OpenAI-compatible API
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.iteracompute.com/v1",
    api_key=os.environ["ITERACOMPUTE_API_KEY"],
)
response = client.chat.completions.create(
    model="qwen/qwen3.8-2.4t-a95b",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

Overview

Text input, streaming and structured outputs; up to 131,072 output tokens.

Pricing

Prices in USD per 1M tokens.

Input$1.95
Cached Input$0.245
Output$5.70

API

Base URL
https://api.iteracompute.com/v1
Endpoint
POST /v1/chat/completions
Model ID
qwen/qwen3.8-2.4t-a95b
Authentication
Authorization: Bearer ITERACOMPUTE_API_KEY

View Docs

Use Cases

Coding, professional work, research and long-horizon agentic tasks.

FAQ

How do I use Qwen3.8-2.4T-A95B with the OpenAI SDK?

Set base_url to https://api.iteracompute.com/v1 and model to qwen/qwen3.8-2.4t-a95b. Authenticate with your IteraCompute API key.

What does Qwen3.8-2.4T-A95B cost?

Input $1.95, cached input $0.245, output $5.70 per 1M tokens, in USD.

What context length is available?

1,048,576 tokens.