← All models
Online

DeepSeek V4.1 Flash

DeepSeek · OpenAI-compatible API

deepseek/deepseek-v4.1-flash

USD / 1M tokens

Input$0.29
Cached Input$0.0058
Output$1.16
Context
1M
Tool Calling
✓
Reasoning
✓
Vision
✗
Streaming
✓
OpenAI-compatible API
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.iteracompute.com/v1",
    api_key=os.environ["ITERACOMPUTE_API_KEY"],
)
response = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

Overview

Text input, streaming and structured outputs; up to 393,216 output tokens.

Pricing

Prices in USD per 1M tokens.

Input$0.29
Cached Input$0.0058
Output$1.16

API

Base URL
https://api.iteracompute.com/v1
Endpoint
POST /v1/chat/completions
Model ID
deepseek/deepseek-v4.1-flash
Authentication
Authorization: Bearer ITERACOMPUTE_API_KEY

View Docs

Use Cases

Efficient agentic coding and million-token context workloads.

FAQ

How do I use DeepSeek V4.1 Flash with the OpenAI SDK?

Set base_url to https://api.iteracompute.com/v1 and model to deepseek/deepseek-v4.1-flash. Authenticate with your IteraCompute API key.

What does DeepSeek V4.1 Flash cost?

Input $0.29, cached input $0.0058, output $1.16 per 1M tokens, in USD.

What context length is available?

1,048,576 tokens.