Skip to content

Models

Explore AI models available on the 0G network

MiniMax-M3
55% OFF

MiniMax-M3, a natively-multimodal model on MiniMax Sparse Attention (MSA); agentic coding, native tool use, and long-horizon tasks. 1M context, thinking on by default.

Context1,000,000
Input
$0.6000$0.27001.3899 0G
Output
$2.4000$1.08005.5700 0G
0GM-1.0-35B-A3B
In-house50% OFFPrivate

0G.AI in-house model optimized for agentic coding and tool use; thinking enabled by default.

Context262,144
Input
$0.1600$0.08000.4120 0G
Output
$0.9600$0.48002.4700 0G
ByteDance Seedance 2.5
28% OFF

ByteDance Seedance 2.5 text-to-video and image-to-video (single first-frame reference via input_reference) — the two Seedance capabilities that map onto a real OpenAI Video API field. ByteDance's own model also supports first+last-frame control and multimodal reference generation (multiple reference images/videos/audio composited per the prompt, including audio-only input), but OpenAI's Video API has no field to express either one, so this integration does not expose a client-facing input for them. Duration 4-30s (default 5); seconds is required on this endpoint and is clamped into that range rather than rejected (the vendor's own 'model-chosen' -1 duration is not reachable through this endpoint — seconds must be a positive number here), resolution 480p/720p/1080p (narrower than earlier Seedance versions — 4k is not served, and is downgraded rather than rejected: a 4K pixel size renders at 1080p, while the bare token '4k' is unrecognised and falls to the 720p default), ratio 16:9/9:16/4:3/3:4/1:1/21:9/adaptive, optional synchronized audio (generate_audio, on by default), fps fixed at 24, optional output_format passthrough (default mp4). Billed on the vendor-reported token count (usage.completion_tokens) — not a flat per-second rate. Async: POST /v1/videos, poll GET /v1/videos/{id} until completed, then GET /v1/videos/{id}/content for the MP4.

Context
Input
Output
$11.7000from$8.4240/1M tok43.4583 0G/1M tokfrom ≈ $0.1028 / sec
Claude Fable 5
10% OFF

Anthropic Claude Fable 5; text and image input, text output, with a 1M-token context window. Extended thinking and tool use supported.

Context1,000,000
Input
$10.0000$9.000046.2100 0G
Output
$50.0000$45.0000231.0900 0G
Qwen3.7-Max
60% OFF

Alibaba Qwen3.7-Max with native function calling and web search; 1M context.

Context1,000,000
Input
$0.82504.2500 0G
Output
$2.475412.7800 0G
GLM-5.2
30% OFF

Zhipu AI GLM-5.2, open-source, purpose-built for long-horizon tasks; 1M lossless context. Strong coding and engineering: autonomous task decomposition, architecture design, full-stack development, integration testing, and multi-platform deployment.

Context1,000,000
Input
$0.96804.9900 0G
Output
$3.388817.4899 0G
DeepSeek-V4-Pro-0813
15% OFF

DeepSeek-V4-Pro for agentic coding, multi-step workflows, and complex reasoning; 1M context, up to 384K output. Currently served as the pinned 2026-08-13 snapshot.

Context1,000,000
Input
$1.27206.5600 0G
Output
$3.816019.7000 0G
DeepSeek-V4-Flash-0731
12% OFF

Lightweight MoE (284B total / 13B active) tuned for fast, low-cost, high-throughput text work; function calling, web search, and thinking on by default (disable with enable_thinking:false); 1M context, up to 384K output. Currently served as the pinned 2026-07-31 snapshot.

Context1,000,000
Input
$0.13790.7120 0G
Output
$0.27501.4100 0G
0GM-1.0-35B-A3B-SIA
In-housePrivate

A 35B hybrid MoE model enhanced with per-token reward guidance. At each decoding step, a 4B Value Model scores candidate tokens and steers generation toward higher-quality outputs, with improvements in harmlessness, helpfulness, and honesty.

Context32,768
Input
$0.53602.7500 0G
Output
$3.216016.5500 0G
GLM-5.3
Private

Zhipu AI GLM-5.3 for long-horizon reasoning, coding, and agentic tool use; 1M context, up to 131K output. Deep thinking is always on and cannot be disabled; how reasoning depth is controlled depends on the provider (top-level reasoning_effort or chat_template_kwargs). Function calling, JSON mode (response_format: json_object), and implicit prompt caching supported.

Context1,048,576
Input
$1.40007.2600 0G
Output
$4.400022.8200 0G
Whisper Large v3
Private

Multilingual automatic speech recognition (ASR); transcription and translation.

Context448
Input
$0.00012999/sec0.00067437 0G/sec
Output
Z-Image-Turbo
Private

Asynchronous text-to-image model with Base64 output. Generates at most 2 images per request — requesting more (n > 2) returns 2 images, not an error.

Context2,048
Input
Output
$0.0088/image0.0456 0G/image
Claude Opus 4.8

Anthropic Claude Opus 4.8; text and image input, text output, with a 1M-token context window. Extended thinking and tool use supported.

Context1,000,000
Input
$5.000025.6700 0G
Output
$25.0000128.3800 0G
Claude Opus 5

Anthropic Claude Opus 5; text and image input, text output, with a 1M-token context window. Extended thinking and tool use supported.

Context1,000,000
Input
$5.000025.6700 0G
Output
$25.0000128.3800 0G
Claude Sonnet 5

Anthropic Claude Sonnet 5; text and image input, text output, with a 1M-token context window. Extended thinking and tool use supported.

Context1,000,000
Input
$1.90009.7500 0G
Output
$9.500048.7800 0G
GLM-5

Zhipu AI GLM-5, purpose-built for coding and agent workflows; 744B foundation, 200K context.

Context202,752
Input
$0.75003.8700 0G
Output
$2.400012.3900 0G
GLM-5.1

Zhipu AI GLM-5.1, purpose-built for long-horizon tasks; 744B foundation, 200K context.

Context206,848
Input
$1.82009.3900 0G
Output
$5.720029.5300 0G
GLM-5.3-Flash

Zhipu AI GLM-5.3-Flash, a fast, cost-effective model for coding and agentic tasks; served via Tencent Cloud MaaS (TokenHub), which exposes both OpenAI and Anthropic faces. Text in / text out, 1M context, up to 128K output. Deep thinking is always on, controllable via reasoning_effort ("none" disables it); function calling (tool_choice limited to auto/none in thinking mode), JSON mode (response_format: json_object), and implicit prompt caching supported.

Context1,000,000
Input
$0.11110.5729 0G
Output
$0.38902.0000 0G
GPT-5.5

GPT-5.5, designed for complex professional workloads, with strong reasoning, high reliability, and improved token efficiency on hard tasks. Text and image input, text output, 1M-token context.

Context1,000,000
Input
$5.000025.6700 0G
Output
$30.0000154.0600 0G
GPT-5.6 Luna

Fast, cost-efficient model in the GPT-5.6 family, optimized for high-volume, cost-sensitive workloads: responsive chat, classification, extraction, lightweight coding, and agentic workflows at lower latency and cost. Text and image input, text output, 1M-token context.

Context1,000,000
Input
$0.19991.0200 0G
Output
$1.20006.1600 0G
GPT-5.6 Sol

Flagship of the GPT-5.6 series, built for advanced reasoning, complex coding, and agentic workflows: multi-step software engineering, long-horizon problem solving, and autonomous tool use. Text and image input, text output, 1M-token context.

Context1,000,000
Input
$5.000025.6700 0G
Output
$30.0000154.0600 0G
GPT-5.6 Terra

Balanced model in the GPT-5.6 family, tuned for workloads that need strong reasoning, coding, and agentic capability at lower cost than the flagship tier. Text and image input, text output, 1M-token context.

Context1,000,000
Input
$2.000010.2700 0G
Output
$12.000061.6200 0G
Hunyuan-3

Tencent Hunyuan 3 (hy3); 295B total / 21B active MoE, native 256K context. Text in / text out. Function calling and implicit prompt caching supported. Served via the Tencent Cloud MaaS (TokenHub international) OpenAI-compatible gateway.

Context262,144
Input
$0.13190.6810 0G
Output
$0.52792.7200 0G
Hunyuan-4-preview

Tencent Hunyuan Hy4 preview; 770B total / 49B active MoE tuned for agent, coding and production workflows — stronger task decomposition, long-horizon tool use and long-chain execution than Hunyuan 3. Served via Tencent Cloud MaaS (TokenHub), which exposes both OpenAI and Anthropic faces. Text in / text out, 1M context, up to 64K output. Deep thinking is on by default and can be disabled with reasoning_effort:"none" (enable_thinking is not honored); function calling, JSON mode (response_format: json_object / json_schema) and implicit prompt caching supported.

Context1,000,000
Input
$0.83404.3000 0G
Output
$2.501012.9100 0G
Kimi-K2.7-Code

Moonshot AI coding model for agentic coding and tool use; multimodal input (text, image, video), thinking always on. 256K context.

Context262,144
Input
$1.23506.3700 0G
Output
$5.200026.8500 0G
Kimi-K3

Moonshot AI Kimi K3; multimodal input (text, image, video), text output. 1M context, deep thinking always on. Function calling and implicit prompt caching supported.

Context1,000,000
Input
$3.000015.4000 0G
Output
$15.000077.0300 0G
MiniMax-H3

MiniMax-H3 (Hailuo-03) text-to-video and image-to-video, billed per generated second. `seconds` must be 4-15, and outside that range it is silently clamped rather than rejected. `size` names a resolution TIER, not output dimensions: send one of the names listed under Resolution Tiers. OpenAI pixel dimensions (1280x720) are also accepted, but they set only the aspect ratio and only for text-to-video — on image-to-video the ratio follows your reference image, so pixel dimensions have no effect there, while a tier name still selects the tier. Prompt required, up to 7000 characters. Image-to-video takes one first-frame image via multipart `input_reference`. Async: POST /v1/videos, poll GET /v1/videos/{id} until completed, then GET /v1/videos/{id}/content for the MP4.

Context
Input
Output
$0.1950/sec1.0044 0G/sec
Qwen3-VL-30B-A3B-Instruct

Alibaba multimodal vision-language model; strong at visual reasoning, OCR, and document understanding.

Context262,144
Input
$0.03580.1850 0G
Output
$0.35871.8500 0G
Qwen3.6-Plus

Alibaba Qwen3.6-Plus with hybrid linear attention and sparse MoE; 1M context, 119 languages.

Context1,000,000
Input
$0.65003.3500 0G
Output
$3.900020.1300 0G
Qwen3.7-Plus

Alibaba multimodal model with vision and video understanding; native function calling, 1M context.

Context1,000,000
Input
$0.52002.6800 0G
Output
$2.080010.7400 0G
Qwen3.8-Flash

Alibaba Qwen3.8-Flash, a fast, cost-effective multimodal model for coding and agentic work; visual and video understanding, function calling, web search, deep thinking, and implicit prompt caching. 1M context.

Context1,000,000
Input
$0.11300.5830 0G
Output
$0.38201.9700 0G
Qwen3.8-Max

Alibaba Qwen3.8-Max, a 2.4T-parameter MoE model for coding and agentic work; native visual understanding of documents and video, function calling and web search, 1M context.

Context1,000,000
Input
$1.65008.5100 0G
Output
$4.950925.5600 0G