← All models
Online

DeepSeek V4 Flash 0731

DeepSeek · OpenAI-compatible API

deepseek/deepseek-v4-flash-0731

USD / 1M tokens

Input$0.29
Cached Input$0.0058
Output$1.16
Context
1M
Tool Calling
✓
Reasoning
✓
Vision
✗
Streaming
✓
OpenAI-compatible API
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.iteracompute.com/v1",
    api_key=os.environ["ITERACOMPUTE_API_KEY"],
)
response = client.chat.completions.create(
    model="deepseek/deepseek-v4-flash-0731",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

Overview

Text input, streaming and structured outputs; up to 393,216 output tokens.

Legacy compatibility alias with the same pricing as DeepSeek V4.1 Flash.

Pricing

Prices in USD per 1M tokens.

Input$0.29
Cached Input$0.0058
Output$1.16

API

Base URL
https://api.iteracompute.com/v1
Endpoint
POST /v1/chat/completions
Model ID
deepseek/deepseek-v4-flash-0731
Authentication
Authorization: Bearer ITERACOMPUTE_API_KEY

View Docs

Use Cases

Efficient agentic coding and long-context workloads.

FAQ

How do I use DeepSeek V4 Flash 0731 with the OpenAI SDK?

Set base_url to https://api.iteracompute.com/v1 and model to deepseek/deepseek-v4-flash-0731. Authenticate with your IteraCompute API key.

What does DeepSeek V4 Flash 0731 cost?

Input $0.29, cached input $0.0058, output $1.16 per 1M tokens, in USD.

What context length is available?

1,048,576 tokens.