← All models
Online

GLM-5.3-Flash

Z.ai · OpenAI-compatible API

z-ai/glm-5.3-flash

USD / 1M tokens

Input$0.099
Cached Input$0.025
Output$0.35
Context
1M
Tool Calling
✓
Reasoning
✓
Vision
✓
Streaming
✓
OpenAI-compatible API
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.iteracompute.com/v1",
    api_key=os.environ["ITERACOMPUTE_API_KEY"],
)
response = client.chat.completions.create(
    model="z-ai/glm-5.3-flash",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

Overview

Text and image input, streaming, reasoning, tool calling and structured outputs; up to 131,072 output tokens.

Pricing

Prices in USD per 1M tokens.

Input$0.099
Cached Input$0.025
Output$0.35

API

Base URL
https://api.iteracompute.com/v1
Endpoint
POST /v1/chat/completions
Model ID
z-ai/glm-5.3-flash
Authentication
Authorization: Bearer ITERACOMPUTE_API_KEY

View Docs

Use Cases

Low-latency coding, agent workflows, long-context analysis and visual understanding.

FAQ

How do I use GLM-5.3-Flash with the OpenAI SDK?

Set base_url to https://api.iteracompute.com/v1 and model to z-ai/glm-5.3-flash. Authenticate with your IteraCompute API key.

What does GLM-5.3-Flash cost?

Input $0.099, cached input $0.025, output $0.35 per 1M tokens, in USD.

What context length is available?

1,048,576 tokens.