1. Create an API key
Create or sign in to your IteraCompute account, then create a key in the Console.
Route requests through one interface.
Create or sign in to your IteraCompute account, then create a key in the Console.
Set api_base and use openai/ before the API Model ID. This prefix is LiteLLM routing metadata.
Use the exact API Model ID for your chosen model. Check its context and capabilities before enabling agent or vision features.
Keep API keys in your server environment or tool credential store. Replace the example key field with your own key where applicable.
import os
from litellm import completion
response = completion(
model="openai/qwen/qwen3.8-27b",
api_base="https://api.iteracompute.com/v1",
api_key=os.environ["ITERACOMPUTE_API_KEY"],
messages=[{"role": "user", "content": "Hello!"}],
max_tokens=128,
)
print(response.choices[0].message.content)model_list:
- model_name: itera-model
litellm_params:
model: openai/qwen/qwen3.8-27b
api_base: https://api.iteracompute.com/v1
api_key: os.environ/ITERACOMPUTE_API_KEY# Bash / zsh / WSL; set ITERACOMPUTE_API_KEY in your environment first.
curl https://api.iteracompute.com/v1/chat/completions \
-H "Authorization: Bearer $ITERACOMPUTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen/qwen3.8-27b","messages":[{"role":"user","content":"Hi"}],"max_tokens":64}'This sends one short request capped at 64 output tokens. Reasoning models may use that budget before returning visible text; increase it if needed. This checks API connectivity, then send a prompt in your tool to check its configuration.
These are API capabilities of the selected model. Feature availability in your workflow also depends on the tool and its settings; reasoning controls vary by model.
Check the key, account balance and model permissions. Restart the tool after changing environment variables.
Use the exact Model ID and keep /v1 in the base URL. Select Chat Completions when the tool offers a Responses API option.
Match context, image support, tool calls and reasoning parameters to the selected model. Aider may need explicit model metadata or edit-format settings for an unfamiliar model.
Reduce concurrency or retry with backoff. For predictable production capacity, explore Reserved Inference.
Connect your key, choose a model and send your first request.