gpt-oss-120b
Open weights (Apache 2.0). OpenRouter from ~$0.04/$0.17 per 1M (provider-dependent). No first-party OpenAI pricing listed.
- Context window
- 131,072 tokens
- Max output
- 131,072 tokens
- Knowledge cutoff
- 2024-06
- Input
- text
- Output
- text
- License
- apache-2.0
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | — | huggingface.co/openai/gpt-oss-120b | — |
| OpenAI API (docs) | gpt-oss-120b | https://api.openai.com/v1/responses | docs |
| Azure OpenAI (Microsoft Foundry) | gpt-oss-120b | — | docs |
| OpenRouter | openai/gpt-oss-120b | openrouter.ai/openai/gpt-oss-120b | — |
Notable capabilities (2)
- Single-GPU open MoE: 117B total / 5.1B active MoE with MXFP4 weights; runs on one 80GB H100/MI300X. source
- Open reasoning with full CoT: Configurable low/medium/high reasoning with full chain-of-thought access, harmony format. source
Self-hosted reasoning/agents via vLLM, Transformers, Ollama, LM Studio.
vllm serve openai/gpt-oss-120b
Sources:
Other OpenAI models
GPT-6 Luna · GPT-6 Sol · GPT Image 2.5 Flare · GPT Image 2.5 Sunburst · GPT-6 Astra · GPT-Live-Transcribe · GPT-Transcribe · GPT-5.6 Terra · GPT-Live 1 · GPT-Realtime-2.1 · GPT-Realtime-2 · GPT-Realtime-Translate · GPT-Rosalind · GPT-Audio-1.5 (and gpt-audio / gpt-audio-mini) · GPT-5.3-Codex · gpt-oss-20b · GPT-4o mini TTS · text-embedding-3-large · text-embedding-3-small · GPT-5.6 Luna · GPT-5.6 Sol · GPT-Realtime-Whisper · GPT-5.5 Pro · GPT-5.5 · GPT Image 2 · GPT-5.4 · GPT-Realtime-1.5 · GPT-4.1 · GPT-4o · TTS-1 / TTS-1 HD · Whisper large-v3 / large-v3-turbo (open weights) · GPT-Realtime and GPT-Realtime mini · o3 · GPT-4o Transcribe / Mini Transcribe / Transcribe Diarize · Whisper (whisper-1 API) · Sora 2