The Elyxir API provides access to hundreds of AI models from major providers. This reference lists the currently active models with Elyxir pricing.
OpenAI Models
GPT-5 Series
Model
Context
Best For
gpt-5.5
128K
Current flagship
gpt-5.5-pro
128K
Maximum reasoning depth
gpt-5.5-instant
128K
Speed-optimized flagship
gpt-5.4
128K
Mid-tier flagship
gpt-5.2
128K
Previous flagship
gpt-5.1
128K
Budget tier
gpt-5-mini
128K
Balanced, cost-effective
gpt-5-nano
128K
Fast, budget-friendly
GPT-4.1 Series
Model
Context
Best For
gpt-4.1
1M
Long context, vision
gpt-4.1-mini
1M
Efficient, vision
gpt-4.1-nano
1M
Ultra-fast, cheap
GPT-4 Series
Model
Context
Best For
gpt-4o
128K
General, Vision
gpt-4o-mini
128K
Cost-effective
Embeddings
Model
Dimensions
Max Input
text-embedding-3-large
3072
8191
text-embedding-3-small
1536
8191
Anthropic Models
Claude 4.8 Series (Latest)
Model
Context
Best For
claude-opus-4.8
200K
Current flagship, agentic coding
Claude 4.7 Series
Model
Context
Best For
claude-opus-4.7
200K
Complex agentic tasks
Claude 4.6 Series
Model
Context
Best For
claude-sonnet-4.6
200K
Coding, Agents
claude-opus-4.6
200K
Most capable
Claude 4.5 Series
Model
Context
Best For
claude-sonnet-4.5
200K
Coding, Agents
claude-opus-4.5
200K
Most capable
claude-haiku-4.5
200K
Fast, affordable
Google Models
Gemini 3.5 Series (Latest)
Model
Context
Best For
gemini-3.5-flash
1M
Fast, beats 3.1 Pro on coding
Gemini 3.1 Series
Model
Context
Best For
gemini-3.1-pro-preview
1M
Most capable, reasoning
Gemini 3 Series
Model
Context
Best For
gemini-3-pro-preview
1M
Advanced reasoning
gemini-3-flash
1M
Fast, efficient
Gemini 2.5 Series
Model
Context
Best For
gemini-2.5-pro
1M
Balanced, long context
gemini-2.5-flash
1M
Fast
gemini-2.5-flash-lite
1M
Ultra-fast, low-cost
xAI Models
Grok 4.3 Series (Latest)
Model
Context
Best For
grok-4.3
1M
Current flagship, cost-efficient
Grok 4 Series
Model
Context
Best For
grok-4.6
1M
High-capability reasoning
grok-4.5
1M
Previous-generation reasoning
grok-4.20
1M
Agentic, multi-step work
grok-4.20-multi-agent
1M
Parallel multi-agent runs
grok-4.20-non-reasoning
1M
Fast, no reasoning pass
grok-build-0.1
1M
Code generation
Mistral AI Models
Model
Context
Best For
mistral-large-3
256K
Most capable
mistral-medium-3.1
128K
Balanced
mistral-small-4
256K
Multimodal, agentic coding
mistral-small-3.2
128K
Fast, efficient
codestral-22b
256K
Code generation
DeepSeek Models
Model
Context
Best For
deepseek-v4-pro
128K
Flagship reasoning and coding
deepseek-v4-flash
128K
Fast, cost-effective V4
deepseek-v3.2
128K
General purpose
deepseek-v3.1
128K
Fast, versatile
Cohere Models
Model
Context
Best For
command-a-plus
256K
Latest flagship, vision
command-a
256K
Enterprise RAG
Perplexity Models
Model
Context
Best For
sonar-deep-research
128K
Deep research with citations
sonar-reasoning-pro
128K
Reasoning with web search
perplexity-sonar-large
128K
General search
perplexity-sonar
128K
Fast search
Meta Models
Llama 4 Series
Model
Context
Best For
llama4-scout
320K
Efficient, vision
llama4-maverick
1M
Complex tasks, vision
Llama 3.3 Series
Model
Context
Best For
llama-3.3-70b
128K
High quality
Open-Weight Models (GLM, Kimi, MiniMax, Qwen)
GLM Series (Zhipu AI)
Model
Context
Best For
glm-5.1
202K
Coding, agentic workflows
glm-5
202K
General purpose
glm-5v-turbo
203K
Native multimodal (text + image)
glm-4.7
128K
Balanced
glm-4.6
202K
Cost-efficient
Kimi Series (Moonshot AI)
Model
Context
Best For
kimi-k2.6
256K
Multimodal, long-horizon coding
kimi-k2.5
256K
Agentic, top-ranked coding
kimi-k2
128K
Agentic workflows
kimi-k2-thinking
128K
Extended reasoning
MiniMax Series
Model
Context
Best For
minimax-m3
1M
Multimodal flagship
minimax-m2.7
200K
Balanced MoE
minimax-m2.5
196K
Cost-effective
Qwen Series (Alibaba)
Model
Context
Best For
qwen3.7-max
1M
Long-horizon agent tasks
qwen3.5
256K
Large-scale multilingual
qwen3-coder-next
256K
Code generation
qwen3-235b-thinking
128K
Extended reasoning
qwen3-32b
40K
Mid-size, versatile
NVIDIA Nemotron Series
Model
Context
Best For
nemotron-3-super-120b
1M
Flagship agentic reasoning
nemotron-3-nano-30b
256K
Fast, low-cost reasoning
Capability-Specific Models
Agentic Coding
Model
Context
Best For
grok-build-0.1
256K
xAI coding model (public beta)
gpt-5.3-codex
128K
OpenAI Responses API coding
Image Generation
Model
Best For
gpt-image-2
Next-gen OpenAI image generation
gpt-image-1.5
Fast OpenAI image generation
nano-banana
Gemini 2.5 Flash Image, precise text
nano-banana-pro
Gemini 3 Pro Image, studio quality
imagen-4
Google Imagen 4 Standard
imagen-4-fast-tg
Imagen 4 fast tier
imagen-4-ultra-tg
Imagen 4 ultra tier
flux-2-pro
FLUX.2 [pro], SOTA open image
flux-2-max
FLUX.2 max tier
flux-2-flex
FLUX.2 flex tier
flux-dev
FLUX.1 dev, high quality
flux-schnell
FLUX.1 schnell, ultra-fast
qwen-image-2.0
Alibaba Qwen-Image 2.0
Video Generation
Model
Best For
sora-2
OpenAI Sora 2
sora-2-pro
OpenAI Sora 2 Pro
The Veo models (veo-2, veo-3, veo-3-fast, veo-3.1, veo-3.1-fast) are no
longer available. Google retired Veo 2.0 and 3.0, and the Veo 3.1 tiers are not
currently served. Sora is the video generation route.
Audio
Model
Best For
eleven-v3
ElevenLabs v3 TTS, natural voice
tts-1-hd
OpenAI HD text-to-speech
tts-1
OpenAI standard text-to-speech
gpt-4o-mini-tts
OpenAI mini text-to-speech
scribe-v2
ElevenLabs speech-to-text (primary)
scribe-v1
ElevenLabs speech-to-text (legacy)
gpt-transcribe
OpenAI transcription (current)
gpt-live-transcribe
Low-latency streaming transcription
gpt-4o-transcribe
GPT-4o powered transcription
gpt-4o-mini-transcribe
Cheaper GPT-4o transcription
whisper-1
OpenAI Whisper transcription
whisper-1, gpt-4o-transcribe and gpt-4o-mini-transcribe are scheduled for
retirement on 2027-02-26. gpt-transcribe (file transcription) and
gpt-live-transcribe (realtime streaming) are the replacements.
Search
Model
Best For
sonar-deep-research
Deep multi-step research
sonar-reasoning-pro
Reasoning with real-time sources
perplexity-sonar-large
Web-grounded answers
perplexity-sonar
Fast web search
Model Selection Guide
By Task
Task
Recommended
General chat
gpt-5-mini, claude-haiku-4.5
Complex reasoning
gpt-5.5, claude-opus-4.8, deepseek-v4-pro
Code generation
claude-sonnet-4.6, grok-build-0.1
Long documents
gpt-4.1, gemini-2.5-pro
Web search
sonar-deep-research, perplexity-sonar-large
Embeddings
text-embedding-3-small
Fast inference
gpt-5-nano, gemini-2.5-flash-lite
Image generation
gpt-image-2, nano-banana-pro
Video generation
sora-2, sora-2-pro
Text-to-speech
eleven-v3, tts-1-hd
Using Model IDs
Use the full model ID in API requests:
response = client.chat.completions.create( model="gpt-4o-mini", # Model ID messages=[...])
Use the ID exactly as listed above — the gateway matches it literally. Two
things that look reasonable but fail:
A provider prefix is not accepted.anthropic/claude-sonnet-4.5 returns a
model-not-found; the ID is claude-sonnet-4.5. (The only IDs containing a
slash are the self-hosted ollama/* models, where the slash is part of the
name itself.)
Versions use dots, not dashes.claude-sonnet-4.6 works;
claude-sonnet-4-6 does not.
Model Availability
Model availability may vary. Check the Model Hub for current availability and pricing.