Choose a model. Know its tradeoffs.

Search the AIVAX catalog snapshot from this build before a workload commits to a provider, context window or capability set. Every row below comes from the public catalog during this build.

Available models
182
Provider groups
31
Fetched
Sep 12, 2026
Two ways to run

Integrated models use the AIVAX balance. BYOK keeps the provider contract and model bill with you.

Available models.

Compare speed, intelligence and token pricing. Models are ordered by release date, newest first; expand an item for provider-specific rates and limits.

Open the source JSON
Showing 182 models

Available AIVAX models ordered by release date, with capabilities, context, token pricing and lifecycle status

  1. deepseek

    deepseek-v4.1-flash

    NewStable
    @deepseek/deepseek-v4.1-flash

    DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the cost-efficient tier of the V4.1 family.

    All contexts
    Input $0.3Output $1.2Cached $0.006
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Tool Calling
    Providers & technical details 12 for @deepseek/deepseek-v4.1-flash

    DeepSeek

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    384,000 tokens
    Throughput
    136 tokens/s
    Latency
    1.112 s
    Uptime
    99.9981%
    Data collection
    Prompt With Training
    All contexts
    Input $0.15Output $0.6Cached $0.003

    DeepInfra:offline

    Disableddeepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    24 tokens/s
    Latency
    3.146 s
    Uptime
    91.9323%
    Data collection
    Zero
    All contexts
    Input $0.2Output $0.6Cached $0.006

    Fireworks

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    105 tokens/s
    Latency
    1.318 s
    Uptime
    99.5117%
    Data collection
    Zero
    All contexts
    Input $0.22Output $0.66Cached $0.007

    Morph

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    27 tokens/s
    Latency
    1.915 s
    Uptime
    99.473%
    Data collection
    Zero
    All contexts
    Input $0.225Output $0.9Cached $0.0225

    SiliconFlow

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    166 tokens/s
    Latency
    1.117 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Modal

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    76 tokens/s
    Latency
    1.3945 s
    Uptime
    95.4147%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.03

    Wafer

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    37 tokens/s
    Latency
    0.9055 s
    Uptime
    99.6663%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Parasail

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    92 tokens/s
    Latency
    1.186 s
    Uptime
    97.7013%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    GMICloud

    deepseek/deepseek-v4.1-flash
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    93 tokens/s
    Latency
    3.4645 s
    Uptime
    99.9774%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Io Net

    deepseek/deepseek-v4.1-flash
    Context
    262,124 tokens
    Max output
    131,072 tokens
    Throughput
    71 tokens/s
    Latency
    1.148 s
    Uptime
    99.8894%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.003

    Novita

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    130 tokens/s
    Latency
    2.059 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Venice

    deepseek/deepseek-v4.1-flash
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    97 tokens/s
    Latency
    0.979 s
    Uptime
    99.5052%
    Data collection
    Zero
    All contexts
    Input $0.375Output $1.5Cached $0.0075
  2. inception

    mercury-2.5

    NewStable
    @inception/mercury-2.5

    Mercury 2.5 is Inception's fastest diffusion reasoning LLM, producing and refining multiple tokens in parallel for agentic and coding workloads.

    All contexts
    Input $0.04Output $0.15Cached $0.004
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    260K
    ThinkingTool CallingDiffusion
    Providers & technical details 1 for @inception/mercury-2.5

    Inception

    inception/mercury-2.5
    Context
    260,000 tokens
    Max output
    65,536 tokens
    Throughput
    76 tokens/s
    Latency
    0.677 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.04Output $0.15Cached $0.004
  3. openai

    gpt-6-astra

    NewStable
    @openai/gpt-6-astra

    GPT-6 Astra is a frontier reasoning model from OpenAI, suited for complex reasoning, coding, and agentic workflows.

    All contexts
    Input $10Output $50Cached $1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Highest
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 5 for @openai/gpt-6-astra

    OpenAI (Flex)

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    58 tokens/s
    Latency
    3.2165 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Azure

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    23 tokens/s
    Latency
    10.211 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $10Output $50Cached $1

    OpenAI

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    33 tokens/s
    Latency
    2.641 s
    Uptime
    99.5477%
    Data collection
    Prompt No Training
    All contexts
    Input $10Output $50Cached $1

    Azure (US)

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    23 tokens/s
    Latency
    4.8155 s
    Data collection
    Zero
    All contexts
    Input $11Output $55Cached $1.1

    OpenAI (Fast)

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    57 tokens/s
    Latency
    3.992 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $20Output $100Cached $2
  4. openai

    gpt-6-astra:pro

    NewStable
    @openai/gpt-6-astra:pro

    GPT-6 Astra is a frontier reasoning model from OpenAI, suited for complex reasoning, coding, and agentic workflows.

    All contexts
    Input $10Output $50Cached $1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 5 for @openai/gpt-6-astra:pro

    OpenAI (Flex)

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    58 tokens/s
    Latency
    3.2165 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Azure

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    23 tokens/s
    Latency
    10.211 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $10Output $50Cached $1

    OpenAI

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    33 tokens/s
    Latency
    2.641 s
    Uptime
    99.5477%
    Data collection
    Prompt No Training
    All contexts
    Input $10Output $50Cached $1

    Azure (US)

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    23 tokens/s
    Latency
    4.8155 s
    Data collection
    Zero
    All contexts
    Input $11Output $55Cached $1.1

    OpenAI (Fast)

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    57 tokens/s
    Latency
    3.992 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $20Output $100Cached $2
  5. google

    gemini-3.8-flash

    NewStable
    @google/gemini-3.8-flash

    Gemini 3.8 Flash is Google's most intelligent Flash model, with significant gains over 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 6 for @google/gemini-3.8-flash

    Google AI Studio (Flex)

    google/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    174 tokens/s
    Latency
    1.3775 s
    Uptime
    99.8152%
    Data collection
    Prompt No Training
    All contexts
    Input $0.375Output $1.875Cached $0.0375Audio input $0.375

    Google (Flex):offline

    Disabledgoogle/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    37 tokens/s
    Latency
    12.23 s
    Uptime
    94.7605%
    Data collection
    Zero
    All contexts
    Input $0.375Output $1.875Cached $0.0375Audio input $0.375

    Google AI Studio

    google/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    187 tokens/s
    Latency
    1.393 s
    Uptime
    99.8256%
    Data collection
    Prompt No Training
    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75

    Google

    google/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    78 tokens/s
    Latency
    2.2695 s
    Uptime
    98.5903%
    Data collection
    Zero
    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75

    Google AI Studio (Priority)

    google/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    28 tokens/s
    Latency
    1.7825 s
    Uptime
    98.4197%
    Data collection
    Prompt No Training
    All contexts
    Input $1.35Output $6.75Cached $0.135Audio input $1.35

    Google (Priority)

    google/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    20 tokens/s
    Latency
    2.0645 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.35Output $6.75Cached $0.135Audio input $1.35
  6. meta

    muse-spark-1.3

    NewStable
    @meta/muse-spark-1.3

    Muse Spark 1.3 is Meta's multimodal reasoning model for long-running agentic, multi-agent, and coding workflows, with emphasis on tracking information across extended tasks and concise execution.

    All contexts
    Input $1.25Output $4.25Cached $0.15
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 2 for @meta/muse-spark-1.3

    Meta (Alt route)

    muse-spark-1.3
    Context
    1,000,000 tokens
    Max output
    1,000,000 tokens
    Data collection
    Prompt With Training
    All contexts
    Input $1.25Output $4.25Cached $0.15

    Meta

    meta/muse-spark-1.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    105 tokens/s
    Latency
    2.2415 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $4.25Cached $0.15
  7. meta

    muse-spark-1.3-contributor

    NewStable
    @meta/muse-spark-1.3-contributor

    Muse Spark 1.3 Contributor is Meta's cost-efficient multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows.

    All contexts
    Input $0.1Output $0.2Cached $0.002
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 2 for @meta/muse-spark-1.3-contributor

    Meta (Alt route)

    muse-spark-1.3-contributor
    Context
    1,000,000 tokens
    Max output
    1,000,000 tokens
    Data collection
    Prompt With Training
    All contexts
    Input $0.1Output $0.2Cached $0.002

    Meta

    meta/muse-spark-1.3-contributor
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    97 tokens/s
    Latency
    2.826 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1Output $0.2Cached $0.002
  8. anthropic

    claude-fable-5.1

    NewStable
    @anthropic/claude-fable-5.1

    Claude Fable 5.1 improves on Fable 5 for agentic coding, long-running workflows, knowledge work, large refactors, visual code generation, finance, and analysis.

    All contexts
    Input $10Output $50Cached $0.25
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 4 for @anthropic/claude-fable-5.1

    Azure

    anthropic/claude-fable-5.1
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    46 tokens/s
    Latency
    5.055 s
    Data collection
    Unknown
    All contexts
    Input $10Output $50Cached $0.25

    Anthropic

    anthropic/claude-fable-5.1
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    33 tokens/s
    Latency
    3.203 s
    Uptime
    99.8412%
    Data collection
    Prompt No Training
    All contexts
    Input $10Output $50Cached $0.25

    Amazon Bedrock

    anthropic/claude-fable-5.1
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $10Output $50Cached $0.25

    Google

    anthropic/claude-fable-5.1
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    49 tokens/s
    Latency
    4.933 s
    Uptime
    100%
    Data collection
    Unknown
    All contexts
    Input $10Output $50Cached $0.25
  9. ibm-granite

    granite-4.2-8b

    NewStable
    @ibm-granite/granite-4.2-8b

    Granite 4.2 8B is IBM's dense reasoning model for mathematics, code generation, multilingual dialogue, and agentic workflows requiring multi-step reasoning.

    All contexts
    Input $0.08Output $0.2Cached $0.0325
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    Medium
    Context up to
    131.1K
    ThinkingTool Calling
    Providers & technical details 2 for @ibm-granite/granite-4.2-8b

    DeepInfra

    ibm-granite/granite-4.2-8b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    87 tokens/s
    Latency
    0.7155 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.06Output $0.25Cached $0.015

    CoreWeave

    ibm-granite/granite-4.2-8b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    97 tokens/s
    Latency
    0.184 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.15Cached $0.05
  10. tencent

    hy4-preview

    PreviewStable
    @tencent/hy4-preview

    Hy4 preview is a 770B-parameter Mixture-of-Experts model from Tencent, with 49B active parameters, designed for coding agents, complex tool-use workflows, and productivity tasks that require planning, context continuity, and sustained multi-step execution.

    All contexts
    Input $0.834Output $2.501Cached $0.042
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    1M
    ThinkingTool Calling
    Providers & technical details 1 for @tencent/hy4-preview

    Tencent

    tencent/hy4-preview
    Context
    1,048,576 tokens
    Max output
    64,000 tokens
    Throughput
    42 tokens/s
    Latency
    3.3265 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.834Output $2.501Cached $0.042
  11. qwen

    qwen3.8-flash

    Stable
    @qwen/qwen3.8-flash

    Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

    All contexts
    Input $0.15Output $0.47Cached $0.016
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    1M
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 2 for @qwen/qwen3.8-flash

    Makora

    qwen/qwen3.8-flash
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    92 tokens/s
    Latency
    0.942 s
    Uptime
    99.7969%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.47Cached $0.016

    Alibaba

    qwen/qwen3.8-flash
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    51 tokens/s
    Latency
    1.556 s
    Uptime
    99.5586%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $0.47Cached $0.016
  12. z-ai

    glm-5.3-flash

    Stable
    @z-ai/glm-5.3-flash

    GLM-5.3-Flash is Z.ai's efficient native multimodal model for coding and long-horizon agent tasks, with image and video understanding and a 1M-token context window.

    All contexts
    Input $0.15Output $0.5Cached $0.03
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1.3M
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 26 for @z-ai/glm-5.3-flash

    DeepInfra

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    18 tokens/s
    Latency
    1.9565 s
    Uptime
    99.1382%
    Data collection
    Zero
    All contexts
    Input $0.075Output $0.25Cached $0.015

    Relace

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    35 tokens/s
    Latency
    1.26 s
    Uptime
    99.9482%
    Data collection
    Zero
    All contexts
    Input $0.09Output $0.3Cached $0.018

    Morph

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.35Cached $0.02

    Wafer

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    9 tokens/s
    Latency
    0.871 s
    Uptime
    99.9709%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.35Cached $0.02

    StreamLake

    z-ai/glm-5.3-flash
    Context
    1,024,000 tokens
    Max output
    128,000 tokens
    Throughput
    43 tokens/s
    Latency
    1.529 s
    Uptime
    99.447%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1124Output $0.3745Cached $0.0225

    GMICloud

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    26 tokens/s
    Latency
    2.083 s
    Uptime
    99.5128%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1125Output $0.375Cached $0.0225

    Reka

    z-ai/glm-5.3-flash
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    14 tokens/s
    Latency
    1.608 s
    Uptime
    99.8756%
    Data collection
    Zero
    All contexts
    Input $0.132Output $0.44Cached $0.0264

    Novita

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    32 tokens/s
    Latency
    3.292 s
    Uptime
    99.7808%
    Data collection
    Zero
    All contexts
    Input $0.132Output $0.44Cached $0.0264

    Makora

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    69 tokens/s
    Latency
    0.831 s
    Uptime
    99.4227%
    Data collection
    Zero
    All contexts
    Input $0.14Output $0.47Cached $0.024

    Crusoe:offline

    Disabledz-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    86 tokens/s
    Latency
    0.876 s
    Uptime
    93.681%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    CoreWeave

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    31 tokens/s
    Latency
    1.239 s
    Uptime
    99.3789%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.05

    Sail Research

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    24 tokens/s
    Latency
    1.899 s
    Uptime
    99.2481%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Fireworks

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    78 tokens/s
    Latency
    0.953 s
    Uptime
    99.8622%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Phala

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    33 tokens/s
    Latency
    3.249 s
    Uptime
    99.7407%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Friendli

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    70 tokens/s
    Latency
    0.7515 s
    Uptime
    95.8315%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $0.5Cached $0.03

    SiliconFlow

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    28 tokens/s
    Latency
    1.656 s
    Uptime
    99.8555%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    DigitalOcean

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    17 tokens/s
    Latency
    2.613 s
    Uptime
    97.3703%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Together

    z-ai/glm-5.3-flash
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    33.5 tokens/s
    Latency
    0.572 s
    Uptime
    99.906%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Parasail

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    60 tokens/s
    Latency
    1.219 s
    Uptime
    99.6513%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    BaseTen

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    135 tokens/s
    Latency
    0.7615 s
    Uptime
    99.9464%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Venice:offline

    Disabledz-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    20 tokens/s
    Latency
    2.583 s
    Uptime
    93.5544%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Io Net

    z-ai/glm-5.3-flash
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    30 tokens/s
    Latency
    1.24 s
    Uptime
    99.8962%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Cloudflare

    z-ai/glm-5.3-flash
    Context
    1,310,720 tokens
    Max output
    1,179,648 tokens
    Throughput
    49 tokens/s
    Latency
    0.8855 s
    Uptime
    99.9447%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Z.AI

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    53 tokens/s
    Latency
    2.516 s
    Uptime
    99.1417%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    NextBit

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    128,000 tokens
    Throughput
    41 tokens/s
    Latency
    2.4155 s
    Uptime
    99.8763%
    Data collection
    Zero
    All contexts
    Input $0.177Output $0.59Cached $0.036

    Modal

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    74 tokens/s
    Latency
    0.6995 s
    Uptime
    99.797%
    Data collection
    Zero
    All contexts
    Input $0.45Output $1.5Cached $0.09
  13. meta

    muse-spark-1.2-contributor

    Stable
    @meta/muse-spark-1.2-contributor

    Muse Spark 1.2 Contributor is Meta's lower-cost reasoning model for experimentation, learning, and early-stage agentic and coding workflows.

    All contexts
    Input $0.1Output $0.2Cached $0.002
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 2 for @meta/muse-spark-1.2-contributor

    Meta (Alt route)

    muse-spark-1.2-contributor
    Context
    1,000,000 tokens
    Max output
    1,000,000 tokens
    Data collection
    Prompt With Training
    All contexts
    Input $0.1Output $0.2Cached $0.002

    Meta

    meta/muse-spark-1.2-contributor
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    85 tokens/s
    Latency
    3.893 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1Output $0.2Cached $0.002
  14. z-ai

    glm-5.3

    Stable
    @z-ai/glm-5.3

    GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window.

    All contexts
    Input $1.4Output $4.4Cached $0.26
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1.3M
    ThinkingTool Calling
    Providers & technical details 26 for @z-ai/glm-5.3

    Inceptron

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    38 tokens/s
    Latency
    0.611 s
    Uptime
    99.2912%
    Data collection
    Zero
    All contexts
    Input $0.8727Output $3.36Cached $0.1639

    Reka

    z-ai/glm-5.3
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    58 tokens/s
    Latency
    1.507 s
    Uptime
    99.5154%
    Data collection
    Zero
    All contexts
    Input $0.936Output $3.168Cached $0.1872

    DigitalOcean

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    43 tokens/s
    Latency
    9.611 s
    Uptime
    97.4692%
    Data collection
    Zero
    All contexts
    Input $0.95Output $3.4Cached $0.2

    Morph

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    79 tokens/s
    Latency
    1.014 s
    Uptime
    98.7516%
    Data collection
    Zero
    All contexts
    Input $1Output $3.41Cached $0.2

    Phala

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    57 tokens/s
    Latency
    1.954 s
    Uptime
    98.5335%
    Data collection
    Zero
    All contexts
    Input $1.05Output $3.3Cached $0.195

    Novita

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    42 tokens/s
    Latency
    2.881 s
    Uptime
    99.929%
    Data collection
    Zero
    All contexts
    Input $1.092Output $3.432Cached $0.2028

    GMICloud

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    48 tokens/s
    Latency
    1.196 s
    Uptime
    99.2525%
    Data collection
    Prompt No Training
    All contexts
    Input $1.12Output $3.52Cached $0.208

    AkashML

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    53.5 tokens/s
    Latency
    4.0715 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.17Output $3.96Cached $0.234

    Decart

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    94 tokens/s
    Latency
    1.072 s
    Uptime
    99.9573%
    Data collection
    Zero
    All contexts
    Input $1.19Output $3.74Cached $0.1955

    Wafer

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    86 tokens/s
    Latency
    2.81 s
    Uptime
    98.6798%
    Data collection
    Zero
    All contexts
    Input $1.19Output $4.4Cached $0.26

    DeepInfra

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    28 tokens/s
    Latency
    9.307 s
    Uptime
    98.4858%
    Data collection
    Zero
    All contexts
    Input $1.2Output $4Cached $0.12

    Sail Research

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    66 tokens/s
    Latency
    1.8615 s
    Uptime
    99.7837%
    Data collection
    Zero
    All contexts
    Input $1.2572Output $3.9512Cached $0.2335

    Friendli

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    125 tokens/s
    Latency
    1.559 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.26Output $3.96Cached $0.234

    Makora

    z-ai/glm-5.3
    Context
    980,000 tokens
    Max output
    128,000 tokens
    Throughput
    111 tokens/s
    Latency
    1.423 s
    Data collection
    Zero
    All contexts
    Input $1.35Output $4.4Cached $0.23

    Crusoe

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    152 tokens/s
    Latency
    0.545 s
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Venice

    z-ai/glm-5.3
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    39.5 tokens/s
    Latency
    4.3025 s
    Uptime
    97.7465%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    SiliconFlow

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    36 tokens/s
    Latency
    1.2375 s
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Together

    z-ai/glm-5.3
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    123 tokens/s
    Latency
    0.591 s
    Uptime
    97.4216%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Parasail

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    60 tokens/s
    Latency
    1.273 s
    Uptime
    99.9025%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Modal

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    88 tokens/s
    Latency
    3.059 s
    Uptime
    99.8686%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    BaseTen:offline

    Disabledz-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    84 tokens/s
    Latency
    1.531 s
    Uptime
    91.8782%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.14

    Fireworks

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    54 tokens/s
    Latency
    1.691 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Cloudflare

    z-ai/glm-5.3
    Context
    1,310,720 tokens
    Max output
    1,179,648 tokens
    Throughput
    50 tokens/s
    Latency
    5.05 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    AtlasCloud

    z-ai/glm-5.3
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    74 tokens/s
    Latency
    4.565 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Z.AI

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    57 tokens/s
    Latency
    2.587 s
    Uptime
    99.6917%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    BaseTen

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    60 tokens/s
    Latency
    0.927 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.1Output $6.6Cached $0.21
  15. qwen

    qwen3.8-27b

    Stable
    @qwen/qwen3.8-27b

    Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interactive agent tasks, with flexible thinking that can be enabled or disabled.

    All contexts
    Input $0.3333Output $2.775Cached $0.05
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    1M
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 15 for @qwen/qwen3.8-27b

    Darkbloom

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    32,768 tokens
    Throughput
    26 tokens/s
    Latency
    1.9755 s
    Uptime
    99.7451%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $2

    DekaLLM

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    63 tokens/s
    Latency
    0.679 s
    Uptime
    97.5967%
    Data collection
    Prompt No Training
    All contexts
    Input $0.2Output $2.5Cached $0.05

    Reka

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    53 tokens/s
    Latency
    0.576 s
    Uptime
    99.9229%
    Data collection
    Zero
    All contexts
    Input $0.214Output $2.55Cached $0.15

    Parasail

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    56 tokens/s
    Latency
    0.801 s
    Uptime
    99.965%
    Data collection
    Zero
    All contexts
    Input $0.24Output $2.2Cached $0.05

    AkashML

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    52 tokens/s
    Latency
    0.52 s
    Uptime
    99.954%
    Data collection
    Zero
    All contexts
    Input $0.25Output $2.2Cached $0.05

    Mancer 2

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    40 tokens/s
    Latency
    0.899 s
    Uptime
    99.7106%
    Data collection
    Zero
    All contexts
    Input $0.25Output $2.75

    Io Net

    qwen/qwen3.8-27b
    Context
    65,536 tokens
    Max output
    58,982 tokens
    Throughput
    24 tokens/s
    Latency
    0.692 s
    Uptime
    98.2318%
    Data collection
    Zero
    All contexts
    Input $0.3Output $2.8Cached $0.18

    Phala

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    46 tokens/s
    Latency
    0.4635 s
    Uptime
    99.2896%
    Data collection
    Zero
    All contexts
    Input $0.3Output $3Cached $0.05

    Chutes

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    31 tokens/s
    Latency
    2.122 s
    Uptime
    99.8926%
    Data collection
    Prompt No Training
    All contexts
    Input $0.32Output $2.5Cached $0.032

    Ionstream

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    47 tokens/s
    Latency
    0.628 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.35Output $2.55Cached $0.05

    CoreWeave

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    61 tokens/s
    Latency
    0.506 s
    Uptime
    99.9358%
    Data collection
    Zero
    All contexts
    Input $0.4Output $3Cached $0.15

    Novita

    qwen/qwen3.8-27b
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    46 tokens/s
    Latency
    2.982 s
    Uptime
    99.7294%
    Data collection
    Zero
    All contexts
    Input $0.42Output $3Cached $0.085

    Alibaba

    qwen/qwen3.8-27b
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    45 tokens/s
    Latency
    0.899 s
    Uptime
    99.8596%
    Data collection
    Prompt No Training
    All contexts
    Input $0.425Output $2.55Cached $0.085

    Cloudflare

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    36 tokens/s
    Latency
    0.5845 s
    Uptime
    98.419%
    Data collection
    Prompt No Training
    All contexts
    Input $0.45Output $3.2Cached $0.05

    Venice

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    70 tokens/s
    Latency
    1.02 s
    Uptime
    99.3289%
    Data collection
    Zero
    All contexts
    Input $0.45Output $3.2
  16. google

    gemini-3.7-flash

    Stable
    @google/gemini-3.7-flash

    Gemini 3.7 Flash is Google's fast multimodal model for agentic workflows, coding, and complex multi-step reasoning.

    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 6 for @google/gemini-3.7-flash

    Google (Flex):offline

    Disabledgoogle/gemini-3.7-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    26 tokens/s
    Latency
    12.496 s
    Uptime
    77.3786%
    Data collection
    Zero
    All contexts
    Input $0.375Output $1.875Cached $0.0375Audio input $0.375

    Google AI Studio (Flex)

    google/gemini-3.7-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    179 tokens/s
    Latency
    1.2405 s
    Uptime
    97.0642%
    Data collection
    Prompt No Training
    All contexts
    Input $0.375Output $1.875Cached $0.0375Audio input $0.375

    Google AI Studio

    google/gemini-3.7-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    192 tokens/s
    Latency
    1.479 s
    Uptime
    98.0058%
    Data collection
    Prompt No Training
    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75

    Google

    google/gemini-3.7-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    78 tokens/s
    Latency
    2.264 s
    Uptime
    99.8739%
    Data collection
    Zero
    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75

    Google (Priority)

    google/gemini-3.7-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    95 tokens/s
    Latency
    2.009 s
    Data collection
    Zero
    All contexts
    Input $1.35Output $6.75Cached $0.135Audio input $1.35

    Google AI Studio (Priority)

    google/gemini-3.7-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    280 tokens/s
    Latency
    2.2005 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.35Output $6.75Cached $0.135Audio input $1.35
  17. deepseek

    deepseek-v4-pro

    Stable
    @deepseek/deepseek-v4-pro

    DeepSeek V4 Pro is a large-scale Mixture-of-Experts model designed for advanced reasoning, coding, and long-horizon agent workflows with 1.6T total parameters.

    All contexts
    Input $1.32Output $3.96Cached $0.044
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Tool Calling
    Providers & technical details 22 for @deepseek/deepseek-v4-pro

    NagaAI:offline

    Disableddeepseek-v4-pro
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.43Output $0.87

    Baidu

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    50 tokens/s
    Latency
    2.2955 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5782Output $1.7345Cached $0.0184

    StreamLake

    deepseek/deepseek-v4-pro-0813
    Context
    1,024,000 tokens
    Max output
    384,000 tokens
    Throughput
    62 tokens/s
    Latency
    3.1595 s
    Uptime
    99.684%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5795Output $1.7384Cached $0.0193

    Alibaba

    deepseek/deepseek-v4-pro-0813
    Context
    1,000,000 tokens
    Max output
    393,216 tokens
    Throughput
    51.5 tokens/s
    Latency
    1.23 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5808Output $1.7424Cached $0.0581

    DeepSeek

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    384,000 tokens
    Throughput
    27 tokens/s
    Latency
    1.228 s
    Uptime
    100%
    Data collection
    Prompt With Training
    All contexts
    Input $0.66Output $1.98Cached $0.022

    Ionstream

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    62 tokens/s
    Latency
    1.123 s
    Uptime
    99.9653%
    Data collection
    Zero
    All contexts
    Input $0.88Output $2.64Cached $0.088

    Novita

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    107 tokens/s
    Latency
    5.455 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.99Output $2.97Cached $0.033

    GMICloud

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    49 tokens/s
    Latency
    5.228 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.056Output $3.168Cached $0.0352

    NextBit

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    27 tokens/s
    Latency
    3.702 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.122Output $3.366Cached $0.037

    DeepInfra

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    16,384 tokens
    Throughput
    98 tokens/s
    Latency
    1.056 s
    Uptime
    99.4444%
    Data collection
    Zero
    All contexts
    Input $1.3Output $2.6Cached $0.1

    CoreWeave

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    116 tokens/s
    Latency
    0.596 s
    Uptime
    99.8523%
    Data collection
    Zero
    All contexts
    Input $1.31Output $3.96Cached $0.044

    Sail Research

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    384,000 tokens
    Throughput
    25 tokens/s
    Latency
    1.194 s
    Uptime
    98.7234%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.044

    BaseTen

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    52 tokens/s
    Latency
    0.537 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.132

    Parasail

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    54 tokens/s
    Latency
    1.1025 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.044

    Together

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    122 tokens/s
    Latency
    0.843 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.13

    DigitalOcean

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    37 tokens/s
    Latency
    0.9145 s
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.044

    SiliconFlow

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    44 tokens/s
    Latency
    1.495 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.044

    BaseTen

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    41 tokens/s
    Latency
    0.461 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.132

    Cloudflare

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    64 tokens/s
    Latency
    1.18 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.32Output $3.96Cached $0.044

    Fireworks

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    63 tokens/s
    Latency
    1.551 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.044

    Phala

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    50 tokens/s
    Latency
    1.583 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.45Output $4.36Cached $0.15

    Venice

    deepseek/deepseek-v4-pro-0813
    Context
    1,000,000 tokens
    Max output
    32,768 tokens
    Throughput
    51 tokens/s
    Latency
    0.954 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.65Output $4.95Cached $0.165
  18. x-ai

    grok-4.6

    Stable
    @x-ai/grok-4.6

    Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

    All contexts
    Input $2Output $6Cached $0.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    500K
    Image InputThinkingTool CallingFile Input
    Providers & technical details 5 for @x-ai/grok-4.6

    xAI (ZDR)

    x-ai/grok-4.6
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    53 tokens/s
    Latency
    0.673 s
    Uptime
    99.7449%
    Data collection
    Zero
    All contexts
    Input $2Output $6Cached $0.5

    xAI

    x-ai/grok-4.6
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    55 tokens/s
    Latency
    1.247 s
    Uptime
    99.8829%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $6Cached $0.5

    Amazon Bedrock (US)

    x-ai/grok-4.6
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Data collection
    Zero
    All contexts
    Input $2.2Output $6.6Cached $0.55

    xAI (Priority) (ZDR)

    x-ai/grok-4.6
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    41 tokens/s
    Latency
    1.84 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $4Output $12Cached $1

    xAI (Priority)

    x-ai/grok-4.6
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    60 tokens/s
    Latency
    1.228 s
    Data collection
    Prompt No Training
    All contexts
    Input $4Output $12Cached $1
  19. nvidia

    nemotron-3.5-lightning

    Stable
    @nvidia/nemotron-3.5-lightning

    NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model with 3B active parameters out of 30B total, designed for high-throughput agentic workloads and specialized tasks.

    All contexts
    Input $0.08Output $0.2Cached $0.04
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    262.1K
    ThinkingTool Calling
    Providers & technical details 4 for @nvidia/nemotron-3.5-lightning

    Darkbloom

    nvidia/nemotron-3.5-lightning
    Context
    262,144 tokens
    Max output
    32,768 tokens
    Throughput
    50 tokens/s
    Latency
    1 s
    Uptime
    96.3303%
    Data collection
    Prompt No Training
    All contexts
    Input $0.065Output $0.18

    Phala

    nvidia/nemotron-3.5-lightning
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    199.5 tokens/s
    Latency
    0.4195 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.08Output $0.2Cached $0.04

    DeepInfra

    nvidia/nemotron-3.5-lightning
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    105 tokens/s
    Latency
    0.333 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.08Output $0.2Cached $0.04

    CoreWeave

    nvidia/nemotron-3.5-lightning
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    283 tokens/s
    Latency
    0.2145 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.25Cached $0.05
  20. upstage

    solar-pro4

    Stable
    @upstage/solar-pro4

    Solar Pro 4 is Upstage's large language model for agentic workflows, office productivity, document-intensive work, and coding.

    All contexts
    Input $0.09Output $0.36Cached $0.018
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    524.3K
    ThinkingTool Calling
    Providers & technical details 2 for @upstage/solar-pro4

    Upstage (ZDR)

    upstage/solar-pro4
    Context
    524,288 tokens
    Max output
    131,072 tokens
    Throughput
    41 tokens/s
    Latency
    2.4065 s
    Uptime
    99.3708%
    Data collection
    Zero
    All contexts
    Input $0.09Output $0.36Cached $0.018

    Upstage

    upstage/solar-pro4
    Context
    524,288 tokens
    Max output
    131,072 tokens
    Throughput
    50 tokens/s
    Latency
    2.06 s
    Uptime
    99.1327%
    Data collection
    Prompt No Training
    All contexts
    Input $0.09Output $0.36Cached $0.018
  21. meta

    muse-glimmer-30b

    Stable
    @meta/muse-glimmer-30b

    Muse Glimmer 30B is Meta's dense, open-weight multimodal model for long-horizon agentic and coding workflows, with image understanding and reliable tool use.

    All contexts
    Input $0.325Output $1.5Cached $0.04
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    131.1K
    Image InputThinkingTool Calling
    Providers & technical details 4 for @meta/muse-glimmer-30b

    Phala

    meta/muse-glimmer-30b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    130 tokens/s
    Latency
    0.685 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.1Cached $0.04

    DeepInfra

    meta/muse-glimmer-30b
    Context
    131,072 tokens
    Max output
    16,384 tokens
    Throughput
    120 tokens/s
    Latency
    0.2055 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.04

    Fireworks

    meta/muse-glimmer-30b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    57 tokens/s
    Latency
    0.0905 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.35Output $1.5Cached $0.04

    Together

    meta/muse-glimmer-30b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    83 tokens/s
    Latency
    0.21 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.35Output $1.5Cached $0.04
  22. meta

    muse-spark-1.2

    Stable
    @meta/muse-spark-1.2

    Muse Spark 1.2 is Meta's multimodal reasoning model for complex agentic tasks, coding, long-horizon workflows, and structured tool use.

    All contexts
    Input $1.25Output $4.25Cached $0.15
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 2 for @meta/muse-spark-1.2

    Meta (Alt route)

    muse-spark-1.2
    Context
    1,000,000 tokens
    Max output
    1,000,000 tokens
    Data collection
    Prompt With Training
    All contexts
    Input $1.25Output $4.25Cached $0.15

    Meta

    meta/muse-spark-1.2
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    95 tokens/s
    Latency
    3.1375 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $4.25Cached $0.15
  23. qwen

    qwen3.8-max

    Stable
    @qwen/qwen3.8-max

    Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding, coding, and agentic workflows.

    All contexts
    Input $2Output $6Cached $0.25
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1M
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 1 for @qwen/qwen3.8-max

    Alibaba

    qwen/qwen3.8-max
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    37 tokens/s
    Latency
    1.769 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $6Cached $0.25
  24. thinkingmachines

    inkling-small

    Stable
    @thinkingmachines/inkling-small

    Inkling Small is Thinking Machines Lab's efficient open-weight multimodal mixture-of-experts model for reasoning, coding, agentic workflows, tool use, and retrieval-augmented generation.

    All contexts
    Input $0.5Output $1.2Cached $0.1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    1M
    Audio InputImage InputThinkingTool Calling
    Providers & technical details 3 for @thinkingmachines/inkling-small

    DeepInfra

    thinkingmachines/inkling-small
    Context
    524,288 tokens
    Max output
    262,144 tokens
    Throughput
    109 tokens/s
    Latency
    0.447 s
    Data collection
    Zero
    All contexts
    Input $0.45Output $1.2Cached $0.1

    BaseTen

    thinkingmachines/inkling-small
    Context
    1,048,576 tokens
    Max output
    32,768 tokens
    Throughput
    26 tokens/s
    Latency
    0.247 s
    Data collection
    Zero
    All contexts
    Input $0.5Output $1.2Cached $0.1

    Together

    thinkingmachines/inkling-small
    Context
    524,288 tokens
    Max output
    471,859 tokens
    Throughput
    106 tokens/s
    Latency
    0.251 s
    Data collection
    Zero
    All contexts
    Input $0.5Output $1.2Cached $0.1
  25. qwen

    qwen3.7-flash

    Stable
    @qwen/qwen3.7-flash

    Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world visual perception.

    All contexts
    Input $0.03Output $0.13Cached $0.006
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    1M
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 1 for @qwen/qwen3.7-flash

    Alibaba

    qwen/qwen3.7-flash
    Context
    1,000,000 tokens
    Max output
    65,536 tokens
    Throughput
    41 tokens/s
    Latency
    0.578 s
    Uptime
    99.996%
    Data collection
    Prompt No Training
    All contexts
    Input $0.03Output $0.13Cached $0.006
  26. anthropic

    claude-5-opus

    Stable
    @anthropic/claude-5-opus

    Claude Opus 5 is Anthropic's flagship model for demanding reasoning, coding, long-horizon agentic work, visual analysis, and complex professional tasks.

    All contexts
    Input $5Output $25Cached $0.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 11 for @anthropic/claude-5-opus

    Azure (US)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    70 tokens/s
    Latency
    2.822 s
    Uptime
    100%
    Data collection
    Unknown
    All contexts
    Input $5Output $25Cached $0.5

    Claude Platform on AWS

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    38 tokens/s
    Latency
    3.96 s
    Uptime
    99.995%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Google

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    70 tokens/s
    Latency
    6.004 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Amazon Bedrock

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    70 tokens/s
    Latency
    3.948 s
    Uptime
    99.9585%
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Azure

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $5Output $25Cached $0.5

    Anthropic

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    67 tokens/s
    Latency
    3.095 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Amazon Bedrock (EU)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Google (US)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Google (EU)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    11 tokens/s
    Latency
    1.473 s
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Amazon Bedrock (US)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    56 tokens/s
    Latency
    4.68 s
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Anthropic (Fast)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    120 tokens/s
    Latency
    2.1815 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $10Output $50Cached $1
  27. google

    gemini-3.5-flash-lite

    Stable
    @google/gemini-3.5-flash-lite

    Gemini 3.5 Flash-Lite is Google's most cost-efficient general availability model, optimized for high-volume agentic tasks, translation, and simple data processing.

    All contexts
    Input $0.315Output $2.625Cached $0.0315Audio input $0.315
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    Medium
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 8 for @google/gemini-3.5-flash-lite

    Google (Flex)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    109 tokens/s
    Latency
    6.608 s
    Uptime
    96.6781%
    Data collection
    Zero
    All contexts
    Input $0.15Output $1.25Cached $0.015Audio input $0.15

    Google AI Studio (Flex)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    167.5 tokens/s
    Latency
    1.9205 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $1.25Cached $0.015Audio input $0.15

    Google AI Studio

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    77 tokens/s
    Latency
    0.512 s
    Uptime
    99.9918%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $2.5Cached $0.03Audio input $0.3

    Google

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    70 tokens/s
    Latency
    0.579 s
    Uptime
    99.9135%
    Data collection
    Zero
    All contexts
    Input $0.3Output $2.5Cached $0.03Audio input $0.3

    Google (EU)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Zero
    All contexts
    Input $0.33Output $2.75Cached $0.033Audio input $0.33

    Google (US)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Zero
    All contexts
    Input $0.33Output $2.75Cached $0.033Audio input $0.33

    Google (Priority)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    10 tokens/s
    Latency
    0.923 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.54Output $4.5Cached $0.054Audio input $0.54

    Google AI Studio (Priority)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    10 tokens/s
    Latency
    0.932 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.54Output $4.5Cached $0.054Audio input $0.54
  28. google

    gemini-3.6-flash

    Stable
    @google/gemini-3.6-flash

    Gemini 3.6 Flash is Google's high-efficiency multimodal model for coding, agentic workflows, and web and app development.

    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 7 for @google/gemini-3.6-flash

    Google (Flex)

    google/gemini-3.6-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    35 tokens/s
    Latency
    6.961 s
    Data collection
    Zero
    All contexts
    Input $0.375Output $1.875Cached $0.0375Audio input $0.375

    Google AI Studio (Flex)

    google/gemini-3.6-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    150 tokens/s
    Latency
    0.846 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.375Output $1.875Cached $0.0375Audio input $0.375

    Google AI Studio

    google/gemini-3.6-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    116 tokens/s
    Latency
    1.707 s
    Uptime
    99.9026%
    Data collection
    Prompt No Training
    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75

    Google

    google/gemini-3.6-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    129 tokens/s
    Latency
    1.414 s
    Uptime
    98.6935%
    Data collection
    Zero
    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75

    Google (Priority)

    google/gemini-3.6-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    179 tokens/s
    Latency
    1.734 s
    Uptime
    99.7753%
    Data collection
    Zero
    All contexts
    Input $1.35Output $6.75Cached $0.135Audio input $1.35

    Google AI Studio (Priority)

    google/gemini-3.6-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    159 tokens/s
    Latency
    1.7635 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.35Output $6.75Cached $0.135Audio input $1.35

    Google (US)

    google/gemini-3.6-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    183 tokens/s
    Latency
    33.478 s
    Data collection
    Zero
    All contexts
    Input $0.825Output $4.125Cached $0.0825Audio input $0.825
  29. poolside

    laguna-s-2.1

    Stable
    @poolside/laguna-s-2.1

    Laguna S 2.1 is Poolside's coding agent model for software engineering and long-horizon agentic workflows, with 118B total parameters and 8B active parameters.

    All contexts
    Input $0.09Output $0.18Cached $0.009
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    ThinkingTool Calling
    Providers & technical details 1 for @poolside/laguna-s-2.1

    Poolside

    poolside/laguna-s-2.1
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    85 tokens/s
    Latency
    0.563 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.09Output $0.18Cached $0.009
  30. meituan

    longcat-2.0

    Stable
    @meituan/longcat-2.0

    LongCat 2.0 is Meituan's sparse mixture-of-experts language model for coding, repository-level changes, long-horizon problem solving, and agentic workflows.

    All contexts
    Input $0.3Output $1.2Cached $0.006
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    ThinkingTool Calling
    Providers & technical details 1 for @meituan/longcat-2.0

    AtlasCloud

    meituan/longcat-2.0
    Context
    1,048,756 tokens
    Max output
    262,144 tokens
    Throughput
    34 tokens/s
    Latency
    2.065 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.006
  31. thinkingmachines

    inkling

    Stable
    @thinkingmachines/inkling

    Inkling is Thinking Machines Lab's open-weight multimodal mixture-of-experts model for general-purpose reasoning, coding, agentic workflows, tool use, and retrieval-augmented generation.

    All contexts
    Input $1Output $4.05Cached $0.17
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    1M
    Audio InputImage InputThinkingTool Calling
    Providers & technical details 4 for @thinkingmachines/inkling

    DeepInfra:offline

    Disabledthinkingmachines/inkling
    Context
    524,288 tokens
    Max output
    262,144 tokens
    Throughput
    143.5 tokens/s
    Latency
    0.7125 s
    Uptime
    94.7598%
    Data collection
    Zero
    All contexts
    Input $0.95Output $4.05Cached $0.16

    BaseTen

    thinkingmachines/inkling
    Context
    1,048,576 tokens
    Max output
    32,768 tokens
    Throughput
    113 tokens/s
    Latency
    0.452 s
    Uptime
    99.3521%
    Data collection
    Zero
    All contexts
    Input $1Output $4.05Cached $0.17

    Together

    thinkingmachines/inkling
    Context
    524,288 tokens
    Max output
    471,859 tokens
    Throughput
    13 tokens/s
    Latency
    8.303 s
    Uptime
    99.0792%
    Data collection
    Zero
    All contexts
    Input $1Output $4.05Cached $0.17

    BaseTen

    thinkingmachines/inkling
    Context
    1,048,576 tokens
    Max output
    32,768 tokens
    Throughput
    113 tokens/s
    Latency
    0.53 s
    Uptime
    98.4314%
    Data collection
    Zero
    All contexts
    Input $1Output $4.05Cached $0.17
  32. moonshotai

    kimi-k3

    Stable
    @moonshotai/kimi-k3

    Kimi K3 is Moonshot AI's ultra-large-scale, open-weight multimodal reasoning model for complex coding, knowledge work, and long-horizon agentic workflows.

    All contexts
    Input $3Output $15Cached $0.3
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool Calling
    Providers & technical details 20 for @moonshotai/kimi-k3

    NagaAI:offline

    Disabledkimi-k3
    Context
    1,048,576 tokens
    Max output
    1,048,576 tokens
    All contexts
    Input $1.5Output $7.5

    Morph

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    11 tokens/s
    Latency
    3.127 s
    Uptime
    99.4364%
    Data collection
    Zero
    All contexts
    Input $2.375Output $13.3Cached $0.2755

    Relace

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    45 tokens/s
    Latency
    1.6025 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.4Output $12Cached $0.24

    Makora:offline

    Disabledmoonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    33 tokens/s
    Latency
    1.071 s
    Uptime
    84.5455%
    Data collection
    Zero
    All contexts
    Input $2.55Output $12.75Cached $0.256

    DigitalOcean

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    20 tokens/s
    Latency
    2.7115 s
    Uptime
    99.5723%
    Data collection
    Zero
    All contexts
    Input $2.55Output $12.95Cached $0.285

    Sail Research

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    71 tokens/s
    Latency
    1.465 s
    Uptime
    99.9739%
    Data collection
    Zero
    All contexts
    Input $2.6481Output $13.2827Cached $0.3026

    Phala

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    33 tokens/s
    Latency
    2.5875 s
    Uptime
    96.3516%
    Data collection
    Zero
    All contexts
    Input $2.85Output $14.25Cached $0.285

    DeepInfra

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    16,384 tokens
    Throughput
    17.5 tokens/s
    Latency
    4.726 s
    Uptime
    99.8598%
    Data collection
    Zero
    All contexts
    Input $2.85Output $14.25Cached $0.285

    Wafer

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    30 tokens/s
    Latency
    0.7305 s
    Uptime
    99.4998%
    Data collection
    Zero
    All contexts
    Input $3Output $12.75Cached $0.3

    Chutes

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    24 tokens/s
    Latency
    2.482 s
    Data collection
    Prompt No Training
    All contexts
    Input $3Output $15Cached $0.3

    Parasail

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    30 tokens/s
    Latency
    1.007 s
    Uptime
    99.7284%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Modal

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    75.5 tokens/s
    Latency
    1.3635 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Together

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    35 tokens/s
    Latency
    1.434 s
    Uptime
    99.7069%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Fireworks

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    42 tokens/s
    Latency
    1.7035 s
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    BaseTen

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    69 tokens/s
    Latency
    1.696 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Moonshot AI

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    24 tokens/s
    Latency
    5.042 s
    Uptime
    99.0595%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Fireworks (US):offline

    Disabledmoonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    33 tokens/s
    Latency
    3.0655 s
    Uptime
    87.6529%
    Data collection
    Zero
    All contexts
    Input $3.3Output $16.5Cached $0.33

    Alibaba

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    35 tokens/s
    Latency
    1.727 s
    Uptime
    99.1111%
    Data collection
    Prompt No Training
    All contexts
    Input $3.45Output $17.25Cached $0.345

    Fireworks (Fast)

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    88 tokens/s
    Latency
    0.6455 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $4.5Output $22.5Cached $0.45

    Morph (Fast)

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    12 tokens/s
    Latency
    2.125 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $6Output $22.5Cached $0.6
  33. kwaipilot

    kat-coder-air-v2.5

    Offline
    @kwaipilot/kat-coder-air-v2.5

    KAT-Coder-Air V2.5 is an efficient agentic coding model designed to autonomously locate, modify, and complete end-to-end software tasks.

    Pricing not published.

    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    Not published
    Tool Calling
    Providers & technical details 0 for @kwaipilot/kat-coder-air-v2.5

    Provider details not published.

  34. kwaipilot

    kat-coder-pro-v2.5

    Stable
    @kwaipilot/kat-coder-pro-v2.5

    KAT-Coder-Pro V2.5 is a flagship agentic coding model designed to autonomously locate, modify, and complete end-to-end software tasks.

    All contexts
    Input $0.74Output $2.96Cached $0.15
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    262.1K
    Tool Calling
    Providers & technical details 1 for @kwaipilot/kat-coder-pro-v2.5

    AtlasCloud

    kwaipilot/kat-coder-pro-v2.5
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $0.74Output $2.96Cached $0.15
  35. openai

    gpt-5.6-luna

    Stable
    @openai/gpt-5.6-luna

    GPT-5.6 Luna is OpenAI's fast, cost-efficient GPT-5.6 model for high-volume chat, classification, and lightweight agentic workflows.

    All contexts
    Input $0.22Output $1.32Cached $0.022
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 7 for @openai/gpt-5.6-luna

    OpenAI (Flex)

    openai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    66 tokens/s
    Latency
    3.831 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1Output $0.6Cached $0.01

    Azure:offline

    Disabledopenai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    68 tokens/s
    Latency
    3.562 s
    Uptime
    85.2456%
    Data collection
    Zero
    All contexts
    Input $0.2Output $1.2Cached $0.02

    OpenAI

    openai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    68 tokens/s
    Latency
    3.372 s
    Uptime
    99.8942%
    Data collection
    Prompt No Training
    All contexts
    Input $0.2Output $1.2Cached $0.02

    Azure (US):offline

    Disabledopenai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    62 tokens/s
    Latency
    4.0455 s
    Uptime
    80.9646%
    Data collection
    Zero
    All contexts
    Input $0.22Output $1.32Cached $0.022

    Amazon Bedrock (US)

    openai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    124 tokens/s
    Latency
    0.543 s
    Uptime
    100%
    Data collection
    Unknown
    All contexts
    Input $0.22Output $1.32Cached $0.022

    Azure (EU)

    openai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    61 tokens/s
    Latency
    1.3555 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.22Output $1.32Cached $0.022

    OpenAI (Fast)

    openai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    87 tokens/s
    Latency
    1.662 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.4Output $2.4Cached $0.04
  36. openai

    gpt-5.6-luna:pro

    Stable
    @openai/gpt-5.6-luna:pro

    GPT-5.6 Luna is OpenAI's fast, cost-efficient GPT-5.6 model for high-volume chat, classification, and lightweight agentic workflows.

    All contexts
    Input $0.2Output $1.2Cached $0.02
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 5 for @openai/gpt-5.6-luna:pro

    OpenAI (Flex)

    openai/gpt-5.6-luna-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    166 tokens/s
    Latency
    28.7865 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1Output $0.6Cached $0.01

    Azure

    openai/gpt-5.6-luna-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    118 tokens/s
    Latency
    14.6795 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.2Output $1.2Cached $0.02

    OpenAI

    openai/gpt-5.6-luna-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    128 tokens/s
    Latency
    9.495 s
    Uptime
    99.8664%
    Data collection
    Prompt No Training
    All contexts
    Input $0.2Output $1.2Cached $0.02

    Azure (EU)

    openai/gpt-5.6-luna-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $0.22Output $1.32Cached $0.022

    OpenAI (Fast)

    openai/gpt-5.6-luna-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    227 tokens/s
    Latency
    6.291 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.4Output $2.4Cached $0.04
  37. openai

    gpt-5.6-sol

    Stable
    @openai/gpt-5.6-sol

    GPT-5.6 Sol is OpenAI's flagship GPT-5.6 model for complex reasoning, coding, and agentic workflows.

    All contexts
    Input $5.5Output $33Cached $0.55
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 7 for @openai/gpt-5.6-sol

    OpenAI (Flex)

    openai/gpt-5.6-sol
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    39 tokens/s
    Latency
    1.5295 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1Output $5Cached $0.1

    OpenAI

    openai/gpt-5.6-sol
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    44 tokens/s
    Latency
    3.9185 s
    Uptime
    99.8166%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $10Cached $0.2

    OpenAI (Fast)

    openai/gpt-5.6-sol
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    35 tokens/s
    Latency
    1.459 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $4Output $20Cached $0.4

    Amazon Bedrock (US)

    openai/gpt-5.6-sol
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $4.4Output $22Cached $0.44

    Azure

    openai/gpt-5.6-sol
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    47 tokens/s
    Latency
    2.8435 s
    Uptime
    99.9742%
    Data collection
    Zero
    All contexts
    Input $5Output $30Cached $0.5

    Azure (US)

    openai/gpt-5.6-sol
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    63 tokens/s
    Latency
    5.389 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5.5Output $33Cached $0.55

    Azure (EU)

    openai/gpt-5.6-sol
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    72 tokens/s
    Latency
    3.333 s
    Data collection
    Zero
    All contexts
    Input $5.5Output $33Cached $0.55
  38. openai

    gpt-5.6-sol:pro

    Stable
    @openai/gpt-5.6-sol:pro

    GPT-5.6 Sol is OpenAI's flagship GPT-5.6 model for complex reasoning, coding, and agentic workflows.

    All contexts
    Input $4.1667Output $24.3333Cached $0.4167
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 5 for @openai/gpt-5.6-sol:pro

    OpenAI (Flex)

    openai/gpt-5.6-sol-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    121 tokens/s
    Latency
    11.4935 s
    Data collection
    Prompt No Training
    All contexts
    Input $1Output $5Cached $0.1

    OpenAI

    openai/gpt-5.6-sol-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    77 tokens/s
    Latency
    13.589 s
    Uptime
    96.5831%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $10Cached $0.2

    OpenAI (Fast)

    openai/gpt-5.6-sol-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    17 tokens/s
    Latency
    1.831 s
    Data collection
    Prompt No Training
    All contexts
    Input $4Output $20Cached $0.4

    Azure

    openai/gpt-5.6-sol-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    77 tokens/s
    Latency
    6.827 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5Output $30Cached $0.5

    Azure (EU)

    openai/gpt-5.6-sol-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $33Cached $0.55
  39. openai

    gpt-5.6-terra

    Stable
    @openai/gpt-5.6-terra

    GPT-5.6 Terra is OpenAI's balanced GPT-5.6 model for everyday coding, reasoning, and agentic tasks.

    All contexts
    Input $2.2Output $13.2Cached $0.22
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 7 for @openai/gpt-5.6-terra

    OpenAI (Flex)

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    82 tokens/s
    Latency
    2.1045 s
    Data collection
    Prompt No Training
    All contexts
    Input $1Output $6Cached $0.1

    Azure

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    54 tokens/s
    Latency
    3.578 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2Output $12Cached $0.2

    OpenAI

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    59 tokens/s
    Latency
    1.9355 s
    Uptime
    98.6473%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $12Cached $0.2

    Azure (US)

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    104 tokens/s
    Latency
    13.961 s
    Data collection
    Zero
    All contexts
    Input $2.2Output $13.2Cached $0.22

    Amazon Bedrock (US)

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $2.2Output $13.2Cached $0.22

    Azure (EU)

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    30 tokens/s
    Latency
    1.581 s
    Data collection
    Zero
    All contexts
    Input $2.2Output $13.2Cached $0.22

    OpenAI (Fast)

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    36 tokens/s
    Latency
    1.2195 s
    Data collection
    Prompt No Training
    All contexts
    Input $4Output $24Cached $0.4
  40. openai

    gpt-5.6-terra:pro

    Stable
    @openai/gpt-5.6-terra:pro

    GPT-5.6 Terra is OpenAI's balanced GPT-5.6 model for everyday coding, reasoning, and agentic tasks.

    All contexts
    Input $2Output $12Cached $0.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 5 for @openai/gpt-5.6-terra:pro

    OpenAI (Flex)

    openai/gpt-5.6-terra-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    338.5 tokens/s
    Latency
    57.4155 s
    Data collection
    Prompt No Training
    All contexts
    Input $1Output $6Cached $0.1

    Azure

    openai/gpt-5.6-terra-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    21 tokens/s
    Latency
    2.2465 s
    Data collection
    Zero
    All contexts
    Input $2Output $12Cached $0.2

    OpenAI

    openai/gpt-5.6-terra-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    106.5 tokens/s
    Latency
    8.2445 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $12Cached $0.2

    Azure (EU)

    openai/gpt-5.6-terra-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $2.2Output $13.2Cached $0.22

    OpenAI (Fast)

    openai/gpt-5.6-terra-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    23 tokens/s
    Latency
    1.911 s
    Data collection
    Prompt No Training
    All contexts
    Input $4Output $24Cached $0.4
  41. x-ai

    grok-4.5

    Stable
    @x-ai/grok-4.5

    Grok 4.5 is xAI's smartest model with frontier performance on coding, knowledge work, and STEM.

    All contexts
    Input $2Output $6Cached $0.3
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    500K
    Image InputThinkingTool CallingFile Input
    Providers & technical details 4 for @x-ai/grok-4.5

    xAI (ZDR)

    x-ai/grok-4.5
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    52 tokens/s
    Latency
    0.722 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2Output $6Cached $0.3

    xAI

    x-ai/grok-4.5
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    50 tokens/s
    Latency
    0.826 s
    Uptime
    99.7176%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $6Cached $0.3

    xAI (Priority) (ZDR)

    x-ai/grok-4.5
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    69 tokens/s
    Latency
    0.736 s
    Data collection
    Zero
    All contexts
    Input $4Output $12Cached $0.6

    xAI (Priority)

    x-ai/grok-4.5
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $4Output $12Cached $0.6
  42. aion-labs

    aion-3.0

    Stable
    @aion-labs/aion-3.0

    Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models.

    All contexts
    Input $3Output $6Cached $0.75
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    131.1K
    ThinkingTool Calling
    Providers & technical details 1 for @aion-labs/aion-3.0

    AionLabs

    aion-labs/aion-3.0
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    57 tokens/s
    Latency
    1.118 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $3Output $6Cached $0.75
  43. aion-labs

    aion-3.0-mini

    Stable
    @aion-labs/aion-3.0-mini

    Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models.

    All contexts
    Input $0.7Output $1.4Cached $0.18
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Medium
    Context up to
    131.1K
    ThinkingTool Calling
    Providers & technical details 1 for @aion-labs/aion-3.0-mini

    AionLabs

    aion-labs/aion-3.0-mini
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    55 tokens/s
    Latency
    0.8195 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.7Output $1.4Cached $0.18
  44. tencent

    hy3

    Stable
    @tencent/hy3

    Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent, built for reasoning, agentic workflows, and real-world production use.

    All contexts
    Input $0.14Output $0.58Cached $0.035
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    ThinkingTool Calling
    Providers & technical details 6 for @tencent/hy3

    DeepInfra:offline

    Disabledtencent/hy3
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    35 tokens/s
    Latency
    0.946 s
    Uptime
    92.4528%
    Data collection
    Zero
    All contexts
    Input $0.07Output $0.29Cached $0.0175

    Tencent

    tencent/hy3
    Context
    262,144 tokens
    Max output
    128,000 tokens
    Throughput
    67 tokens/s
    Latency
    2.22 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.0825Output $0.33Cached $0.0206

    Novita

    tencent/hy3
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    61 tokens/s
    Latency
    2.691 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.14Output $0.58Cached $0.035

    GMICloud

    tencent/hy3
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    12 tokens/s
    Latency
    1.862 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.14Output $0.58Cached $0.035

    Phala

    tencent/hy3
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    27 tokens/s
    Latency
    1.679 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.64Cached $0.04

    AtlasCloud

    tencent/hy3
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    110 tokens/s
    Latency
    1.193 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.2Output $0.8Cached $0.05
  45. poolside

    laguna-xs-2.1

    Stable
    @poolside/laguna-xs-2.1

    Laguna XS 2.1 is Poolside's compact coding agent model in the 33B-A3B category, combining tool calling and reasoning for agentic software engineering tasks.

    All contexts
    Input $0.06Output $0.12Cached $0.03
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    262.1K
    ThinkingTool Calling
    Providers & technical details 1 for @poolside/laguna-xs-2.1

    Poolside

    poolside/laguna-xs-2.1
    Context
    262,144 tokens
    Max output
    32,768 tokens
    Throughput
    158.5 tokens/s
    Latency
    0.269 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.06Output $0.12Cached $0.03
  46. anthropic

    claude-5-sonnet

    Stable
    @anthropic/claude-5-sonnet

    Claude Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work.

    All contexts
    Input $2Output $10Cached $0.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 10 for @anthropic/claude-5-sonnet

    Claude Platform on AWS

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    38 tokens/s
    Latency
    1.8165 s
    Uptime
    99.9962%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $10Cached $0.2

    Azure (US)

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    64 tokens/s
    Latency
    0.637 s
    Uptime
    100%
    Data collection
    Unknown
    All contexts
    Input $2Output $10Cached $0.2

    Azure

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    32 tokens/s
    Latency
    2.717 s
    Data collection
    Unknown
    All contexts
    Input $2Output $10Cached $0.2

    Google

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    61 tokens/s
    Latency
    2.501 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2Output $10Cached $0.2

    Amazon Bedrock

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    51 tokens/s
    Latency
    2.5175 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2Output $10Cached $0.2

    Anthropic

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    75 tokens/s
    Latency
    1.277 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $10Cached $0.2

    Amazon Bedrock (EU)

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    124 tokens/s
    Latency
    4.293 s
    Data collection
    Zero
    All contexts
    Input $2.2Output $11Cached $0.22

    Google (US)

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $2.2Output $11Cached $0.22

    Google (EU)

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    65 tokens/s
    Latency
    3.3815 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.2Output $11Cached $0.22

    Amazon Bedrock (US)

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    68.5 tokens/s
    Latency
    4.339 s
    Data collection
    Zero
    All contexts
    Input $2.2Output $11Cached $0.22
  47. nex-agi

    nex-n2-mini

    Offline
    @nex-agi/nex-n2-mini

    Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI for coding, tool use, deep research, and long-horizon agentic workflows.

    Pricing not published.

    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    Not published
    Image InputThinkingTool Calling
    Providers & technical details 0 for @nex-agi/nex-n2-mini

    Provider details not published.

  48. z-ai

    glm-5.2

    Stable
    @z-ai/glm-5.2

    GLM-5.2 is Z.ai's flagship model for long-horizon tasks, with a usable 1M-token context window for project-level engineering context and long-running agent workflows.

    All contexts
    Input $1.4Output $4.4Cached $0.26
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    ThinkingTool Calling
    Providers & technical details 33 for @z-ai/glm-5.2

    Baidu

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    57 tokens/s
    Latency
    0.919 s
    Uptime
    99.9907%
    Data collection
    Prompt No Training
    All contexts
    Input $0.4802Output $1.5092Cached $0.0892

    StreamLake

    z-ai/glm-5.2
    Context
    1,024,000 tokens
    Max output
    128,000 tokens
    Throughput
    45 tokens/s
    Latency
    1.5895 s
    Uptime
    99.4881%
    Data collection
    Prompt No Training
    All contexts
    Input $0.4805Output $1.5101Cached $0.0892

    DeepInfra

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    163,840 tokens
    Throughput
    46 tokens/s
    Latency
    1.206 s
    Uptime
    99.518%
    Data collection
    Zero
    All contexts
    Input $0.4875Output $1.56Cached $0.091

    Ambient

    z-ai/glm-5.2
    Context
    202,752 tokens
    Max output
    182,476 tokens
    Throughput
    58 tokens/s
    Latency
    2.725 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.6Output $2Cached $0.15

    Baidu (Fast)

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    65 tokens/s
    Latency
    1.507 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.6818Output $2.3876Cached $0.1697

    Novita

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    42 tokens/s
    Latency
    2.4825 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.6832Output $2.1472Cached $0.1269

    DigitalOcean

    z-ai/glm-5.2
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    60 tokens/s
    Latency
    0.8815 s
    Uptime
    98.7235%
    Data collection
    Zero
    All contexts
    Input $0.7Output $2.2Cached $0.105

    CoreWeave

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    87.5 tokens/s
    Latency
    0.6855 s
    Uptime
    99.2857%
    Data collection
    Zero
    All contexts
    Input $0.76Output $2.42Cached $0.14

    AtlasCloud

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    40.5 tokens/s
    Latency
    1.8015 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.938Output $2.948Cached $0.1742

    Alibaba

    z-ai/glm-5.2
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    47 tokens/s
    Latency
    1.189 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.966Output $3.036Cached $0.1932

    Inceptron

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    72 tokens/s
    Latency
    0.4335 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.0833Output $2.8931Cached $0.1764

    SiliconFlow

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    78.5 tokens/s
    Latency
    1.9905 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.19Output $3.74Cached $0.221

    Phala

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    93 tokens/s
    Latency
    1.346 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.26Output $3Cached $0.22

    Mistral (ZDR)

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    128,000 tokens
    Throughput
    70 tokens/s
    Latency
    1.9545 s
    Uptime
    99.9856%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.14

    Mistral

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    128,000 tokens
    Throughput
    77 tokens/s
    Latency
    1.304 s
    Uptime
    99.9508%
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.14

    BaseTen

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    124 tokens/s
    Latency
    1.545 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.14

    Crusoe

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    69 tokens/s
    Latency
    0.993 s
    Uptime
    99.3064%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Together

    z-ai/glm-5.2
    Context
    512,000 tokens
    Max output
    460,800 tokens
    Throughput
    74 tokens/s
    Latency
    0.578 s
    Uptime
    97.693%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Fireworks

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    61 tokens/s
    Latency
    0.982 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.14

    BaseTen

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    104.5 tokens/s
    Latency
    0.828 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.14

    Venice

    z-ai/glm-5.2
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    90 tokens/s
    Latency
    1.986 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    GMICloud

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    39 tokens/s
    Latency
    1.891 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Parasail

    z-ai/glm-5.2
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    104 tokens/s
    Latency
    0.889 s
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Friendli

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    91 tokens/s
    Latency
    0.335 s
    Uptime
    99.9793%
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Cloudflare

    z-ai/glm-5.2
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    44 tokens/s
    Latency
    1.8355 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Z.AI

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    47 tokens/s
    Latency
    3.318 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Mistral (EU)

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    128,000 tokens
    Throughput
    207 tokens/s
    Latency
    0.543 s
    Data collection
    Zero
    All contexts
    Input $1.54Output $4.84Cached $0.154

    Decart (Fast)

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    179 tokens/s
    Latency
    1.0065 s
    Uptime
    99.6598%
    Data collection
    Zero
    All contexts
    Input $2.1Output $6.6Cached $0.21

    BaseTen (Fast)

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    125.5 tokens/s
    Latency
    0.958 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.1Output $6.6Cached $0.21

    Fireworks (Fast)

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    129 tokens/s
    Latency
    0.997 s
    Uptime
    99.3927%
    Data collection
    Zero
    All contexts
    Input $2.1Output $6.6Cached $0.21

    BaseTen (Fast)

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    107.5 tokens/s
    Latency
    2.0385 s
    Data collection
    Zero
    All contexts
    Input $2.1Output $6.6Cached $0.21

    Fireworks (Fast)

    z-ai/glm-5.2
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    119 tokens/s
    Latency
    1.103 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.1Output $6.6Cached $0.21

    Alibaba (Fast)

    z-ai/glm-5.2
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    69 tokens/s
    Latency
    0.9215 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $2.31Output $7.26Cached $0.462
  49. moonshotai

    kimi-k2.7-code

    Stable
    @moonshotai/kimi-k2.7-code

    Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts.

    All contexts
    Input $0.95Output $4Cached $0.19
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 15 for @moonshotai/kimi-k2.7-code

    DeepInfra

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    16,384 tokens
    Throughput
    34 tokens/s
    Latency
    1.03 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.68Output $3.4Cached $0.136

    Inceptron

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    58 tokens/s
    Latency
    0.821 s
    Uptime
    99.9452%
    Data collection
    Zero
    All contexts
    Input $0.7062Output $3.21Cached $0.18

    CoreWeave

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    98.5 tokens/s
    Latency
    1.3935 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.71Output $3.5Cached $0.15

    StreamLake

    moonshotai/kimi-k2.7-code
    Context
    256,000 tokens
    Max output
    32,000 tokens
    Throughput
    67 tokens/s
    Latency
    1.11 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.7125Output $3Cached $0.1425

    Venice

    moonshotai/kimi-k2.7-code
    Context
    256,000 tokens
    Max output
    65,536 tokens
    Throughput
    29.5 tokens/s
    Latency
    1.218 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.75Output $3.5Cached $0.16

    ModelRun

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    137 tokens/s
    Latency
    0.386 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.85Output $3.75Cached $0.16

    SiliconFlow

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    26 tokens/s
    Latency
    1.2315 s
    Data collection
    Zero
    All contexts
    Input $0.8592Output $3.8Cached $0.1799

    Novita

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    61 tokens/s
    Latency
    1.4245 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.912Output $3.84Cached $0.1824

    Fireworks

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    110 tokens/s
    Latency
    0.402 s
    Data collection
    Zero
    All contexts
    Input $0.95Output $4Cached $0.19

    BaseTen

    moonshotai/kimi-k2.7-code
    Context
    262,000 tokens
    Max output
    235,800 tokens
    Throughput
    140 tokens/s
    Latency
    0.4285 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.95Output $4Cached $0.16

    GMICloud

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    60 tokens/s
    Latency
    3.1145 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.95Output $4Cached $0.19

    Alibaba

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    16,384 tokens
    Throughput
    55 tokens/s
    Latency
    2.785 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.95Output $4Cached $0.19

    Cloudflare

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    93 tokens/s
    Latency
    1.202 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.95Output $4Cached $0.19

    Moonshot AI

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    72 tokens/s
    Latency
    1.294 s
    Data collection
    Zero
    All contexts
    Input $0.95Output $4Cached $0.19

    Moonshot AI

    moonshotai/kimi-k2.7-code
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    11 tokens/s
    Latency
    2.631 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.9Output $8Cached $0.38
  50. anthropic

    claude-fable-5

    Stable
    @anthropic/claude-fable-5

    Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window.

    All contexts
    Input $10Output $50Cached $1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 7 for @anthropic/claude-fable-5

    NagaAI:offline

    Disabledclaude-fable-5
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    All contexts
    Input $5Output $25

    Claude Platform on AWS

    anthropic/claude-fable-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    43 tokens/s
    Latency
    4.339 s
    Data collection
    Prompt No Training
    All contexts
    Input $10Output $50Cached $1

    Azure

    anthropic/claude-fable-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    25 tokens/s
    Latency
    4.204 s
    Data collection
    Unknown
    All contexts
    Input $10Output $50Cached $1

    Amazon Bedrock

    anthropic/claude-fable-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $10Output $50Cached $1

    Google

    anthropic/claude-fable-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    50 tokens/s
    Latency
    7.5325 s
    Data collection
    Unknown
    All contexts
    Input $10Output $50Cached $1

    Anthropic

    anthropic/claude-fable-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    53 tokens/s
    Latency
    4.5075 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $10Output $50Cached $1

    Google (EU)

    anthropic/claude-fable-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $11Output $55Cached $1.1
  51. nex-agi

    nex-n2-pro

    Offline
    @nex-agi/nex-n2-pro

    Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total, built for coding, tool use, and long-horizon agentic workflows.

    Pricing not published.

    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    Not published
    Image InputThinkingTool Calling
    Providers & technical details 0 for @nex-agi/nex-n2-pro

    Provider details not published.

  52. nvidia

    nemotron-3-ultra

    Stable
    @nvidia/nemotron-3-ultra

    NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture.

    All contexts
    Input $0.6Output $2.4Cached $0.12
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    ThinkingTool Calling
    Providers & technical details 4 for @nvidia/nemotron-3-ultra

    DeepInfra

    nvidia/nemotron-3-ultra-550b-a55b
    Context
    262,144 tokens
    Max output
    16,384 tokens
    Throughput
    232 tokens/s
    Latency
    2.288 s
    Data collection
    Zero
    All contexts
    Input $0.5Output $2.2Cached $0.1

    BaseTen

    nvidia/nemotron-3-ultra-550b-a55b
    Context
    202,800 tokens
    Max output
    182,520 tokens
    Throughput
    129 tokens/s
    Latency
    0.268 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.4Cached $0.12

    BaseTen

    nvidia/nemotron-3-ultra-550b-a55b
    Context
    202,800 tokens
    Max output
    182,520 tokens
    Throughput
    123 tokens/s
    Latency
    0.258 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.4Cached $0.12

    Venice

    nvidia/nemotron-3-ultra-550b-a55b
    Context
    256,000 tokens
    Max output
    32,768 tokens
    Throughput
    28 tokens/s
    Latency
    1.347 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.625Output $3.125Cached $0.1875
  53. qwen

    qwen3.7-plus

    Stable
    @qwen/qwen3.7-plus

    Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its vision-language abilities.

    All contexts
    Input $0.32Output $1.28Cached $0.064
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Image InputTool CallingFile Input
    Providers & technical details 1 for @qwen/qwen3.7-plus

    Alibaba

    qwen/qwen3.7-plus
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    12 tokens/s
    Latency
    0.776 s
    Uptime
    99.9978%
    Data collection
    Prompt No Training
    All contexts
    Input $0.32Output $1.28Cached $0.064
  54. minimax

    m3

    Stable
    @minimax/m3

    MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use.

    All contexts
    Input $0.3Output $1.2Cached $0.06
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 13 for @minimax/m3

    NagaAI:offline

    Disabledminimax-m3
    Context
    524,288 tokens
    Max output
    524,288 tokens
    All contexts
    Input $0.15Output $0.6

    CoreWeave

    minimax/minimax-m3
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    46 tokens/s
    Latency
    0.576 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.23Output $0.96Cached $0.05

    GMICloud

    minimax/minimax-m3
    Context
    1,048,576 tokens
    Max output
    524,288 tokens
    Throughput
    57 tokens/s
    Latency
    1.066 s
    Uptime
    99.6851%
    Data collection
    Prompt No Training
    All contexts
    Input $0.24Output $0.96Cached $0.048

    DeepInfra

    minimax/minimax-m3
    Context
    524,288 tokens
    Max output
    512,000 tokens
    Throughput
    23.5 tokens/s
    Latency
    2.3585 s
    Uptime
    98.9899%
    Data collection
    Zero
    All contexts
    Input $0.28Output $1.1Cached $0.056

    StreamLake

    minimax/minimax-m3
    Context
    1,000,000 tokens
    Max output
    512,000 tokens
    Throughput
    63 tokens/s
    Latency
    1.333 s
    Uptime
    99.2432%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Venice:offline

    Disabledminimax/minimax-m3
    Context
    524,288 tokens
    Max output
    65,536 tokens
    Throughput
    67 tokens/s
    Latency
    0.89 s
    Uptime
    85.879%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Together

    minimax/minimax-m3
    Context
    524,288 tokens
    Max output
    471,859 tokens
    Throughput
    45 tokens/s
    Latency
    1.416 s
    Uptime
    99.6058%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Parasail

    minimax/minimax-m3
    Context
    1,048,576 tokens
    Max output
    524,288 tokens
    Throughput
    75 tokens/s
    Latency
    0.4835 s
    Uptime
    99.9324%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.06

    AtlasCloud

    minimax/minimax-m3
    Context
    524,300 tokens
    Max output
    524,288 tokens
    Throughput
    110 tokens/s
    Latency
    2.709 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Novita

    minimax/minimax-m3
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    73 tokens/s
    Latency
    1.595 s
    Uptime
    99.8412%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Minimax

    minimax/minimax-m3
    Context
    524,288 tokens
    Max output
    512,000 tokens
    Throughput
    96 tokens/s
    Latency
    0.817 s
    Uptime
    99.1853%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.06

    SambaNova

    minimax/minimax-m3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    136 tokens/s
    Latency
    1.819 s
    Uptime
    99.3007%
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.4

    ModelRun

    minimax/minimax-m3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    126 tokens/s
    Latency
    0.658 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.75Output $3Cached $0.15
  55. anthropic

    claude-4.8-opus

    Stable
    @anthropic/claude-4.8-opus

    Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window.

    All contexts
    Input $5Output $25Cached $0.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 11 for @anthropic/claude-4.8-opus

    Claude Platform on AWS

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    57 tokens/s
    Latency
    1.329 s
    Uptime
    99.9166%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Azure (US)

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    23 tokens/s
    Latency
    2.451 s
    Data collection
    Unknown
    All contexts
    Input $5Output $25Cached $0.5

    Azure

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    20.5 tokens/s
    Latency
    2.49 s
    Data collection
    Unknown
    All contexts
    Input $5Output $25Cached $0.5

    Amazon Bedrock

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    71 tokens/s
    Latency
    3.441 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Google

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    65 tokens/s
    Latency
    1.9115 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Anthropic

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    63 tokens/s
    Latency
    1.633 s
    Uptime
    99.6896%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Amazon Bedrock (US)

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Google (US)

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Amazon Bedrock (EU)

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Google (EU)

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Anthropic (Fast)

    anthropic/claude-opus-4.8
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    102 tokens/s
    Latency
    0.729 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $10Output $50Cached $1
  56. qwen

    qwen3.7-max

    Stable
    @qwen/qwen3.7-max

    Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks, and long-horizon autonomous agents.

    All contexts
    Input $1.475Output $4.425Cached $0.295
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1M
    Tool Calling
    Providers & technical details 1 for @qwen/qwen3.7-max

    Alibaba

    qwen/qwen3.7-max
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    47 tokens/s
    Latency
    0.791 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.475Output $4.425Cached $0.295
  57. x-ai

    grok-build-0.1

    Stable
    @x-ai/grok-build-0.1

    Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding agents, tool use, and multi-step development.

    All contexts
    Input $1Output $2Cached $0.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    256K
    Image InputTool Calling
    Providers & technical details 4 for @x-ai/grok-build-0.1

    xAI

    x-ai/grok-build-0.1
    Context
    256,000 tokens
    Max output
    230,400 tokens
    Throughput
    109 tokens/s
    Latency
    0.992 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1Output $2Cached $0.2

    xAI (ZDR)

    x-ai/grok-build-0.1
    Context
    256,000 tokens
    Max output
    230,400 tokens
    Throughput
    69 tokens/s
    Latency
    0.305 s
    Data collection
    Zero
    All contexts
    Input $1Output $2Cached $0.2

    xAI (Priority) (ZDR)

    x-ai/grok-build-0.1
    Context
    256,000 tokens
    Max output
    230,400 tokens
    Data collection
    Zero
    All contexts
    Input $2Output $4Cached $0.4

    xAI (Priority)

    x-ai/grok-build-0.1
    Context
    256,000 tokens
    Max output
    230,400 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $4Cached $0.4
  58. google

    gemini-3.5-flash

    Stable
    @google/gemini-3.5-flash

    Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed.

    All contexts
    Input $1.5Output $9Cached $0.15Audio input $3
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 8 for @google/gemini-3.5-flash

    NagaAI:offline

    Disabledgemini-3.5-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.75Output $4.5

    Google (Flex)

    google/gemini-3.5-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    6 tokens/s
    Latency
    12.785 s
    Data collection
    Zero
    All contexts
    Input $0.75Output $4.5Cached $0.075Audio input $1.5

    Google AI Studio (Flex)

    google/gemini-3.5-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    115 tokens/s
    Latency
    1.318 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.75Output $4.5Cached $0.075Audio input $1.5

    Google

    google/gemini-3.5-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    116 tokens/s
    Latency
    1.694 s
    Uptime
    99.9171%
    Data collection
    Zero
    All contexts
    Input $1.5Output $9Cached $0.15Audio input $3

    Google AI Studio

    google/gemini-3.5-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    18 tokens/s
    Latency
    1.519 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.5Output $9Cached $0.15Audio input $3

    Google (US)

    google/gemini-3.5-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Zero
    All contexts
    Input $1.65Output $9.9Cached $0.165Audio input $3.3

    Google (Priority)

    google/gemini-3.5-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    86 tokens/s
    Latency
    0.8485 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.7Output $16.2Cached $0.27Audio input $5.4

    Google AI Studio (Priority)

    google/gemini-3.5-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    86 tokens/s
    Latency
    1.1015 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $2.7Output $16.2Cached $0.27Audio input $5.4
  59. google

    gemini-3.1-flash-lite

    PreviewStable
    @google/gemini-3.1-flash-lite

    Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases.

    All contexts
    Input $0.2625Output $1.575Cached $0.0263Audio input $0.525
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    Medium
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 9 for @google/gemini-3.1-flash-lite

    NagaAI:offline

    Disabledgemini-3.1-flash-lite-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.13Output $0.75

    Google (Flex)

    google/gemini-3.1-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    41 tokens/s
    Latency
    15.644 s
    Uptime
    99.3328%
    Data collection
    Zero
    All contexts
    Input $0.125Output $0.75Cached $0.0125Audio input $0.25

    Google AI Studio (Flex)

    google/gemini-3.1-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    18 tokens/s
    Latency
    0.414 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.125Output $0.75Cached $0.0125Audio input $0.25

    Google AI Studio

    google/gemini-3.1-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    84 tokens/s
    Latency
    0.668 s
    Uptime
    99.9693%
    Data collection
    Prompt No Training
    All contexts
    Input $0.25Output $1.5Cached $0.025Audio input $0.5

    Google

    google/gemini-3.1-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    82 tokens/s
    Latency
    0.885 s
    Uptime
    99.9965%
    Data collection
    Zero
    All contexts
    Input $0.25Output $1.5Cached $0.025Audio input $0.5

    Google (EU)

    google/gemini-3.1-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Zero
    All contexts
    Input $0.275Output $1.65Cached $0.0275Audio input $0.55

    Google (US)

    google/gemini-3.1-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    36.5 tokens/s
    Latency
    1.6245 s
    Data collection
    Zero
    All contexts
    Input $0.275Output $1.65Cached $0.0275Audio input $0.55

    Google (Priority)

    google/gemini-3.1-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    36.5 tokens/s
    Latency
    0.618 s
    Data collection
    Zero
    All contexts
    Input $0.45Output $2.7Cached $0.045Audio input $0.9

    Google AI Studio (Priority)

    google/gemini-3.1-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    129 tokens/s
    Latency
    0.449 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.45Output $2.7Cached $0.045Audio input $0.9
  60. x-ai

    grok-4.3

    Stable
    @x-ai/grok-4.3

    Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual logic.

    All contexts
    Input $1.25Output $2.5Cached $0.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 4 for @x-ai/grok-4.3

    xAI (ZDR)

    x-ai/grok-4.3
    Context
    1,000,000 tokens
    Max output
    900,000 tokens
    Throughput
    89 tokens/s
    Latency
    0.591 s
    Uptime
    99.9273%
    Data collection
    Zero
    All contexts
    Input $1.25Output $2.5Cached $0.2

    xAI

    x-ai/grok-4.3
    Context
    1,000,000 tokens
    Max output
    900,000 tokens
    Throughput
    88 tokens/s
    Latency
    0.666 s
    Uptime
    99.9603%
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $2.5Cached $0.2

    xAI (Priority) (ZDR)

    x-ai/grok-4.3
    Context
    1,000,000 tokens
    Max output
    900,000 tokens
    Throughput
    19.5 tokens/s
    Latency
    0.539 s
    Data collection
    Zero
    All contexts
    Input $2.5Output $5Cached $0.4

    xAI (Priority)

    x-ai/grok-4.3
    Context
    1,000,000 tokens
    Max output
    900,000 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $2.5Output $5Cached $0.4
  61. qwen

    qwen3.6-27b

    Stable
    @qwen/qwen3.6-27b

    Qwen3.6 27B is a dense 27-billion-parameter model from Alibaba Qwen, designed for agentic coding, long-context reasoning, and multimodal workflows.

    All contexts
    Input $0.31Output $2.95
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 7 for @qwen/qwen3.6-27b

    Groq (Alt route)

    qwen/qwen3.6-27b
    Context
    131,072 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.6Output $3

    Chutes:offline

    Disabledqwen/qwen3.6-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    24 tokens/s
    Latency
    2.578 s
    Uptime
    94.7619%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $2Cached $0.03

    SiliconFlow

    qwen/qwen3.6-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    29 tokens/s
    Latency
    1.499 s
    Data collection
    Zero
    All contexts
    Input $0.3Output $3.2

    Phala:offline

    Disabledqwen/qwen3.6-27b
    Context
    262,144 tokens
    Max output
    262,140 tokens
    Throughput
    24 tokens/s
    Latency
    1.257 s
    Uptime
    84.8993%
    Data collection
    Zero
    All contexts
    Input $0.32Output $2.7Cached $0.15

    DeepInfra

    qwen/qwen3.6-27b
    Context
    262,144 tokens
    Max output
    81,920 tokens
    Throughput
    39 tokens/s
    Latency
    1.3255 s
    Uptime
    99.7024%
    Data collection
    Zero
    All contexts
    Input $0.32Output $3.2

    Venice

    qwen/qwen3.6-27b
    Context
    256,000 tokens
    Max output
    65,536 tokens
    Throughput
    37 tokens/s
    Latency
    0.783 s
    Data collection
    Zero
    All contexts
    Input $0.325Output $3.25

    Alibaba

    qwen/qwen3.6-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    56 tokens/s
    Latency
    0.836 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.45Output $2.7
  62. deepseek

    deepseek-v4-flash

    Stable
    @deepseek/deepseek-v4-flash

    An efficiency-optimized Mixture-of-Experts model from DeepSeek designed for fast inference and high-throughput workloads.

    All contexts
    Input $0.14Output $0.28Cached $0.028
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Tool Calling
    Providers & technical details 17 for @deepseek/deepseek-v4-flash

    NagaAI:offline

    Disableddeepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    384,000 tokens
    All contexts
    Input $0.04Output $0.08

    StreamLake

    deepseek/deepseek-v4-flash
    Context
    1,024,000 tokens
    Max output
    384,000 tokens
    Throughput
    55 tokens/s
    Latency
    1.584 s
    Uptime
    98.5444%
    Data collection
    Prompt No Training
    All contexts
    Input $0.0657Output $0.1313Cached $0.0131

    Baidu

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    84 tokens/s
    Latency
    0.643 s
    Uptime
    99.9966%
    Data collection
    Prompt No Training
    All contexts
    Input $0.0658Output $0.1316Cached $0.0132

    DigitalOcean

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    9 tokens/s
    Latency
    1.295 s
    Uptime
    99.9528%
    Data collection
    Zero
    All contexts
    Input $0.0679Output $0.168Cached $0.0168

    DeepInfra

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    20 tokens/s
    Latency
    1.111 s
    Uptime
    98.7424%
    Data collection
    Zero
    All contexts
    Input $0.09Output $0.18Cached $0.018

    GMICloud

    deepseek/deepseek-v4-flash
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    54 tokens/s
    Latency
    1.603 s
    Uptime
    99.9786%
    Data collection
    Prompt No Training
    All contexts
    Input $0.091Output $0.182Cached $0.0182

    Venice

    deepseek/deepseek-v4-flash
    Context
    1,000,000 tokens
    Max output
    32,768 tokens
    Throughput
    42 tokens/s
    Latency
    1.099 s
    Uptime
    98.1458%
    Data collection
    Zero
    All contexts
    Input $0.0966Output $0.1925Cached $0.0196

    Wafer

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    91 tokens/s
    Latency
    0.449 s
    Uptime
    99.9757%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.25Cached $0.05

    SiliconFlow

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    67 tokens/s
    Latency
    2.125 s
    Uptime
    99.6113%
    Data collection
    Zero
    All contexts
    Input $0.13Output $0.28Cached $0.028

    Alibaba

    deepseek/deepseek-v4-flash
    Context
    1,000,000 tokens
    Max output
    393,216 tokens
    Throughput
    86 tokens/s
    Latency
    0.8655 s
    Uptime
    99.9491%
    Data collection
    Prompt No Training
    All contexts
    Input $0.134Output $0.268Cached $0.0268

    Novita

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    47 tokens/s
    Latency
    1.276 s
    Uptime
    99.9984%
    Data collection
    Zero
    All contexts
    Input $0.14Output $0.28Cached $0.028

    AtlasCloud

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    37 tokens/s
    Latency
    1.024 s
    Uptime
    99.9935%
    Data collection
    Prompt No Training
    All contexts
    Input $0.14Output $0.28Cached $0.028

    Parasail

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    52 tokens/s
    Latency
    0.733 s
    Uptime
    99.9692%
    Data collection
    Zero
    All contexts
    Input $0.14Output $0.28Cached $0.07

    NextBit

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    64 tokens/s
    Latency
    1.578 s
    Uptime
    99.8441%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.35Cached $0.035

    Mancer 2

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    57 tokens/s
    Latency
    0.686 s
    Uptime
    99.6918%
    Data collection
    Zero
    All contexts
    Input $0.19Output $0.5

    Phala

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    62 tokens/s
    Latency
    1.384 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.2Output $0.4Cached $0.07

    Azure (US)

    deepseek/deepseek-v4-flash
    Context
    1,048,576 tokens
    Max output
    384,000 tokens
    Throughput
    57 tokens/s
    Latency
    0.9425 s
    Uptime
    99.0635%
    Data collection
    Zero
    All contexts
    Input $0.21Output $0.56Cached $0.031
  63. openai

    gpt-5.5

    Stable
    @openai/gpt-5.5

    GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning and higher reliability.

    All contexts
    Input $5.5Output $33Cached $0.55
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Highest
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 8 for @openai/gpt-5.5

    NagaAI:offline

    Disabledgpt-5.5
    Context
    1,050,000 tokens
    Max output
    131,072 tokens
    All contexts
    Input $2.5Output $15

    OpenAI (Flex)

    openai/gpt-5.5
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    16 tokens/s
    Latency
    1.026 s
    Data collection
    Prompt No Training
    All contexts
    Input $2.5Output $15Cached $0.25

    Azure

    openai/gpt-5.5
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    102 tokens/s
    Latency
    2.8625 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5Output $30Cached $0.5

    OpenAI

    openai/gpt-5.5
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    36 tokens/s
    Latency
    1.887 s
    Uptime
    99.9647%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $30Cached $0.5

    Azure (US)

    openai/gpt-5.5
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $33Cached $0.55

    Azure (EU)

    openai/gpt-5.5
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    52 tokens/s
    Latency
    4.8275 s
    Data collection
    Zero
    All contexts
    Input $5.5Output $33Cached $0.55

    Amazon Bedrock (US)

    openai/gpt-5.5
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $5.5Output $33Cached $0.55

    OpenAI (Fast)

    openai/gpt-5.5
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    96 tokens/s
    Latency
    3.326 s
    Data collection
    Prompt No Training
    All contexts
    Input $12.5Output $75Cached $1.25
  64. openai

    gpt-5.5-pro

    Stable
    @openai/gpt-5.5-pro

    GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window.

    All contexts
    Input $22.5Output $135
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 3 for @openai/gpt-5.5-pro

    NagaAI:offline

    Disabledgpt-5.5-pro
    Context
    1,050,000 tokens
    Max output
    131,072 tokens
    All contexts
    Input $15Output $90

    OpenAI (Flex)

    openai/gpt-5.5-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $15Output $90

    OpenAI

    openai/gpt-5.5-pro
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    5 tokens/s
    Latency
    3.406 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $30Output $180
  65. xiaomi

    mimo-v2.5

    Stable
    @xiaomi/mimo-v2.5

    MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks.

    All contexts
    Input $0.168Output $0.336Cached $0.003
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool Calling
    Providers & technical details 6 for @xiaomi/mimo-v2.5

    Xiaomi:offline

    Disabledmimo-v2.5
    Context
    1,050,000 tokens
    Max output
    131,072 tokens

    Pricing not published.

    GMICloud:offline

    Disabledxiaomi/mimo-v2.5
    Context
    1,050,000 tokens
    Max output
    945,000 tokens
    Throughput
    2 tokens/s
    Latency
    4.4005 s
    Uptime
    66.619%
    Data collection
    Prompt No Training
    All contexts
    Input $0.119Output $0.238Cached $0.0026

    DeepInfra

    xiaomi/mimo-v2.5
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    13 tokens/s
    Latency
    1.766 s
    Uptime
    99.8196%
    Data collection
    Zero
    All contexts
    Input $0.133Output $0.266Cached $0.0027

    Xiaomi

    xiaomi/mimo-v2.5
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    41 tokens/s
    Latency
    2.8035 s
    Uptime
    97.1612%
    Data collection
    Prompt No Training
    All contexts
    Input $0.14Output $0.28Cached $0.0028

    StreamLake

    xiaomi/mimo-v2.5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    42 tokens/s
    Latency
    1.368 s
    Uptime
    98.3%
    Data collection
    Prompt No Training
    All contexts
    Input $0.168Output $0.336Cached $0.0034

    Novita

    xiaomi/mimo-v2.5
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    36 tokens/s
    Latency
    3.665 s
    Uptime
    99.3772%
    Data collection
    Zero
    All contexts
    Input $0.168Output $0.336Cached $0.0034
  66. xiaomi

    mimo-v2.5-pro

    Stable
    @xiaomi/mimo-v2.5-pro

    MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro.

    All contexts
    Input $0.435Output $0.87Cached $0.0036
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    1M
    ThinkingTool Calling
    Providers & technical details 8 for @xiaomi/mimo-v2.5-pro

    Xiaomi:offline

    Disabledmimo-v2.5-pro
    Context
    1,050,000 tokens
    Max output
    131,072 tokens

    Pricing not published.

    GMICloud:offline

    Disabledxiaomi/mimo-v2.5-pro
    Context
    1,050,000 tokens
    Max output
    945,000 tokens
    Throughput
    18 tokens/s
    Latency
    4.206 s
    Uptime
    57.7358%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3045Output $0.609Cached $0.0028

    DeepInfra

    xiaomi/mimo-v2.5-pro
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    29 tokens/s
    Latency
    0.361 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.39Output $1.17Cached $0.078

    DigitalOcean

    xiaomi/mimo-v2.5-pro
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    21 tokens/s
    Latency
    0.893 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.4Output $1.5Cached $0.08

    AtlasCloud

    xiaomi/mimo-v2.5-pro
    Context
    1,024,000 tokens
    Max output
    131,072 tokens
    Throughput
    36 tokens/s
    Latency
    4.947 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.435Output $0.87Cached $0.0036

    Xiaomi

    xiaomi/mimo-v2.5-pro
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    27 tokens/s
    Latency
    2.6565 s
    Uptime
    97.866%
    Data collection
    Prompt No Training
    All contexts
    Input $0.435Output $0.87Cached $0.0036

    Novita

    xiaomi/mimo-v2.5-pro
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    28 tokens/s
    Latency
    5.238 s
    Uptime
    97.992%
    Data collection
    Zero
    All contexts
    Input $0.4802Output $0.9605Cached $0.004

    StreamLake:offline

    Disabledxiaomi/mimo-v2.5-pro
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    27 tokens/s
    Latency
    2.106 s
    Uptime
    79.8319%
    Data collection
    Prompt No Training
    All contexts
    Input $0.522Output $1.044Cached $0.0043
  67. moonshotai

    kimi-k2.6

    Stable
    @moonshotai/kimi-k2.6

    Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration.

    All contexts
    Input $0.95Output $4Cached $0.16
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    Image InputTool Calling
    Providers & technical details 21 for @moonshotai/kimi-k2.6

    Baidu

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    39 tokens/s
    Latency
    1.0045 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5795Output $2.44Cached $0.0976

    Chutes

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    65,535 tokens
    Throughput
    6 tokens/s
    Latency
    2.536 s
    Uptime
    99.4264%
    Data collection
    Prompt No Training
    All contexts
    Input $0.58Output $3.4Cached $0.058

    Decart

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    41 tokens/s
    Latency
    0.384 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.5865Output $2.4696Cached $0.0988

    Inceptron

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    79 tokens/s
    Latency
    0.29 s
    Uptime
    99.973%
    Data collection
    Zero
    All contexts
    Input $0.5895Output $2.45Cached $0.17

    StreamLake

    moonshotai/kimi-k2.6
    Context
    256,000 tokens
    Max output
    230,400 tokens
    Throughput
    9 tokens/s
    Latency
    1.99 s
    Uptime
    99.4413%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5985Output $2.52Cached $0.1008

    CoreWeave

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    38 tokens/s
    Latency
    0.4945 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.65Output $3.41Cached $0.15

    Crusoe

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    20 tokens/s
    Latency
    0.234 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.7Output $3.5Cached $0.35

    DeepInfra

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    16,384 tokens
    Throughput
    15 tokens/s
    Latency
    2.1605 s
    Uptime
    99.359%
    Data collection
    Zero
    All contexts
    Input $0.75Output $3.5Cached $0.15

    Venice

    moonshotai/kimi-k2.6
    Context
    256,000 tokens
    Max output
    65,536 tokens
    Throughput
    9 tokens/s
    Latency
    2.072 s
    Uptime
    99.8031%
    Data collection
    Zero
    All contexts
    Input $0.75Output $3.5Cached $0.16

    Parasail

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    42 tokens/s
    Latency
    0.713 s
    Uptime
    99.3428%
    Data collection
    Zero
    All contexts
    Input $0.75Output $3.5Cached $0.16

    SiliconFlow

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    10 tokens/s
    Latency
    2.155 s
    Uptime
    99.8875%
    Data collection
    Zero
    All contexts
    Input $0.77Output $3.4Cached $0.14

    Novita

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    8 tokens/s
    Latency
    2.865 s
    Uptime
    99.8415%
    Data collection
    Zero
    All contexts
    Input $0.8Output $3.4Cached $0.16

    GMICloud

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    40 tokens/s
    Latency
    2.659 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.855Output $3.6Cached $0.144

    DigitalOcean

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    48 tokens/s
    Latency
    0.409 s
    Data collection
    Zero
    All contexts
    Input $0.95Output $4Cached $0.19

    AtlasCloud

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    9 tokens/s
    Latency
    1.488 s
    Uptime
    99.3865%
    Data collection
    Prompt No Training
    All contexts
    Input $0.95Output $4Cached $0.16

    Cloudflare

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    36 tokens/s
    Latency
    0.888 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.95Output $4Cached $0.16

    Moonshot AI

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    33 tokens/s
    Latency
    2.566 s
    Uptime
    99.9563%
    Data collection
    Zero
    All contexts
    Input $0.95Output $4Cached $0.16

    Phala

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    7 tokens/s
    Latency
    2.808 s
    Uptime
    99.6528%
    Data collection
    Zero
    All contexts
    Input $1.09Output $4.6Cached $0.37

    BaseTen

    moonshotai/kimi-k2.6
    Context
    262,000 tokens
    Max output
    235,800 tokens
    Throughput
    125 tokens/s
    Latency
    1.076 s
    Data collection
    Zero
    All contexts
    Input $0.95Output $4Cached $0.16

    Fireworks

    moonshotai/kimi-k2.6
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    36 tokens/s
    Latency
    0.618 s
    Data collection
    Zero
    All contexts
    Input $0.95Output $4Cached $0.16

    BaseTen

    moonshotai/kimi-k2.6
    Context
    262,000 tokens
    Max output
    235,800 tokens
    Throughput
    155 tokens/s
    Latency
    4.012 s
    Data collection
    Zero
    All contexts
    Input $0.95Output $4Cached $0.16
  68. anthropic

    claude-4.7-opus

    Stable
    @anthropic/claude-4.7-opus

    Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents.

    All contexts
    Input $5Output $25Cached $0.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 8 for @anthropic/claude-4.7-opus

    Claude Platform on AWS

    anthropic/claude-opus-4.7
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    41 tokens/s
    Latency
    0.964 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Azure

    anthropic/claude-opus-4.7
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $5Output $25Cached $0.5

    Google

    anthropic/claude-opus-4.7
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    62 tokens/s
    Latency
    1.284 s
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Amazon Bedrock

    anthropic/claude-opus-4.7
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    57 tokens/s
    Latency
    1.72 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Anthropic

    anthropic/claude-opus-4.7
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    55 tokens/s
    Latency
    0.955 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Google (US)

    anthropic/claude-opus-4.7
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Amazon Bedrock (EU)

    anthropic/claude-opus-4.7
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    66 tokens/s
    Latency
    1.768 s
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Google (EU)

    anthropic/claude-opus-4.7
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55
  69. z-ai

    glm-5.1

    Stable
    @z-ai/glm-5.1

    GLM-5.1 represents a major advance in coding ability, with especially notable improvements in tackling long-horizon tasks.

    All contexts
    Input $1.4Output $4.4Cached $0.26
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    204.8K
    ThinkingTool Calling
    Providers & technical details 16 for @z-ai/glm-5.1

    NagaAI:offline

    Disabledglm-5.1
    Context
    204,800 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.63Output $1.98

    Baidu

    z-ai/glm-5.1
    Context
    202,752 tokens
    Max output
    131,072 tokens
    Throughput
    60 tokens/s
    Latency
    0.925 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.9646Output $3.0316Cached $0.1791

    StreamLake

    z-ai/glm-5.1
    Context
    200,000 tokens
    Max output
    128,000 tokens
    Throughput
    40 tokens/s
    Latency
    1.364 s
    Uptime
    99.8525%
    Data collection
    Prompt No Training
    All contexts
    Input $0.966Output $3.036Cached $0.1794

    Chutes:offline

    Disabledz-ai/glm-5.1
    Context
    202,752 tokens
    Max output
    65,535 tokens
    Throughput
    30 tokens/s
    Latency
    5.016 s
    Uptime
    77.4834%
    Data collection
    Prompt No Training
    All contexts
    Input $0.98Output $3.08Cached $0.098

    DeepInfra

    z-ai/glm-5.1
    Context
    202,752 tokens
    Max output
    65,536 tokens
    Throughput
    65 tokens/s
    Latency
    0.847 s
    Uptime
    99.8252%
    Data collection
    Zero
    All contexts
    Input $1.05Output $3.5Cached $0.205

    SiliconFlow

    z-ai/glm-5.1
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    28 tokens/s
    Latency
    1.838 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.19Output $3.74Cached $0.6

    Crusoe

    z-ai/glm-5.1
    Context
    202,752 tokens
    Max output
    182,476 tokens
    Throughput
    70 tokens/s
    Latency
    0.516 s
    Data collection
    Zero
    All contexts
    Input $1.2Output $4.4Cached $0.25

    Phala:offline

    Disabledz-ai/glm-5.1
    Context
    202,752 tokens
    Max output
    128,000 tokens
    Throughput
    30 tokens/s
    Latency
    2.476 s
    Uptime
    87.2464%
    Data collection
    Zero
    All contexts
    Input $1.21Output $4.2Cached $0.6

    AtlasCloud

    z-ai/glm-5.1
    Context
    202,752 tokens
    Max output
    182,476 tokens
    Throughput
    56 tokens/s
    Latency
    1.208 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.26Output $3.96Cached $0.234

    Alibaba

    z-ai/glm-5.1
    Context
    202,745 tokens
    Max output
    131,072 tokens
    Throughput
    60 tokens/s
    Latency
    1.356 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.33Output $4.18Cached $0.247

    Novita

    z-ai/glm-5.1
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    42 tokens/s
    Latency
    2.2115 s
    Data collection
    Zero
    All contexts
    Input $1.38Output $4.4Cached $0.26

    Nebius

    z-ai/glm-5.1
    Context
    202,752 tokens
    Max output
    182,476 tokens
    Throughput
    32 tokens/s
    Latency
    0.9585 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4

    GMICloud

    z-ai/glm-5.1
    Context
    202,752 tokens
    Max output
    182,476 tokens
    Throughput
    46 tokens/s
    Latency
    1.702 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Friendli

    z-ai/glm-5.1
    Context
    202,752 tokens
    Max output
    182,476 tokens
    Throughput
    143 tokens/s
    Latency
    0.089 s
    Uptime
    99.97%
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Z.AI

    z-ai/glm-5.1
    Context
    202,752 tokens
    Max output
    131,072 tokens
    Throughput
    24 tokens/s
    Latency
    9.4735 s
    Uptime
    99.1968%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Venice

    z-ai/glm-5.1
    Context
    200,000 tokens
    Max output
    80,000 tokens
    Throughput
    64 tokens/s
    Latency
    1.145 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.4014Output $4.4044Cached $0.2603
  70. google

    gemma-4-26b-a4b-it

    Stable
    @google/gemma-4-26b-a4b-it

    Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at a fraction of the compute cost. Supports multimodal input including text, images, and video.

    All contexts
    Input $0.13Output $0.4
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    262.1K
    Image InputThinkingTool CallingFile Input
    Providers & technical details 11 for @google/gemma-4-26b-a4b-it

    Darkbloom

    google/gemma-4-26b-a4b-it
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    26 tokens/s
    Latency
    1.124 s
    Uptime
    99.8552%
    Data collection
    Prompt No Training
    All contexts
    Input $0.042Output $0.22

    DekaLLM

    google/gemma-4-26b-a4b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    86 tokens/s
    Latency
    0.475 s
    Uptime
    99.9364%
    Data collection
    Prompt No Training
    All contexts
    Input $0.06Output $0.33

    DeepInfra

    google/gemma-4-26b-a4b-it
    Context
    262,144 tokens
    Max output
    16,384 tokens
    Throughput
    22 tokens/s
    Latency
    0.727 s
    Uptime
    99.556%
    Data collection
    Zero
    All contexts
    Input $0.07Output $0.34

    NextBit

    google/gemma-4-26b-a4b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    56 tokens/s
    Latency
    0.323 s
    Uptime
    99.9766%
    Data collection
    Zero
    All contexts
    Input $0.09Output $0.3Cached $0.05

    Cloudflare

    google/gemma-4-26b-a4b-it
    Context
    256,000 tokens
    Max output
    230,400 tokens
    Throughput
    53 tokens/s
    Latency
    0.408 s
    Uptime
    99.9604%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1Output $0.3

    Makora

    google/gemma-4-26b-a4b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    47 tokens/s
    Latency
    0.676 s
    Uptime
    99.4784%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.34Cached $0.034

    Venice

    google/gemma-4-26b-a4b-it
    Context
    256,000 tokens
    Max output
    8,192 tokens
    Throughput
    16 tokens/s
    Latency
    1.3705 s
    Uptime
    99.0047%
    Data collection
    Zero
    All contexts
    Input $0.13Output $0.4Cached $0.05

    Parasail

    google/gemma-4-26b-a4b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    14 tokens/s
    Latency
    1.077 s
    Uptime
    98.8791%
    Data collection
    Zero
    All contexts
    Input $0.13Output $0.4Cached $0.05

    Novita

    google/gemma-4-26b-a4b-it
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    38 tokens/s
    Latency
    0.8805 s
    Uptime
    99.8204%
    Data collection
    Zero
    All contexts
    Input $0.13Output $0.4

    SiliconFlow

    google/gemma-4-26b-a4b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    28 tokens/s
    Latency
    1.463 s
    Uptime
    99.716%
    Data collection
    Zero
    All contexts
    Input $0.14Output $0.4Cached $0.05

    Google

    google/gemma-4-26b-a4b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    31 tokens/s
    Latency
    1.161 s
    Uptime
    96.6926%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.6
  71. google

    gemma-4-31b-it

    Stable
    @google/gemma-4-31b-it

    Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input. Strong on coding, reasoning, and document understanding tasks with multilingual support across 140+ languages.

    All contexts
    Input $0.14Output $0.4
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    Image InputThinkingTool Calling
    Providers & technical details 14 for @google/gemma-4-31b-it

    DeepInfra

    google/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    16,384 tokens
    Throughput
    34 tokens/s
    Latency
    0.789 s
    Uptime
    98.8604%
    Data collection
    Zero
    All contexts
    Input $0.09Output $0.34Cached $0.05

    CoreWeave

    google/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    44 tokens/s
    Latency
    0.85 s
    Uptime
    96.7144%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.34Cached $0.1

    Venice

    google/gemma-4-31b-it
    Context
    256,000 tokens
    Max output
    8,192 tokens
    Throughput
    20 tokens/s
    Latency
    0.661 s
    Uptime
    99.888%
    Data collection
    Zero
    All contexts
    Input $0.12Output $0.36Cached $0.09

    Chutes

    google/gemma-4-31b-it
    Context
    131,072 tokens
    Max output
    65,536 tokens
    Throughput
    14 tokens/s
    Latency
    3.1595 s
    Uptime
    97.1334%
    Data collection
    Prompt No Training
    All contexts
    Input $0.12Output $0.37Cached $0.012

    DeepInfra

    google/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    16,384 tokens
    Throughput
    26 tokens/s
    Latency
    0.993 s
    Uptime
    97.413%
    Data collection
    Zero
    All contexts
    Input $0.13Output $0.38

    Crusoe:offline

    Disabledgoogle/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    262,141 tokens
    Throughput
    11 tokens/s
    Latency
    1.068 s
    Uptime
    78.9474%
    Data collection
    Zero
    All contexts
    Input $0.14Output $0.4Cached $0.14

    Friendli

    google/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    8,192 tokens
    Throughput
    97 tokens/s
    Latency
    0.5095 s
    Uptime
    99.5772%
    Data collection
    Prompt No Training
    All contexts
    Input $0.14Output $0.4

    Novita:offline

    Disabledgoogle/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    5 tokens/s
    Latency
    1.176 s
    Uptime
    72.7727%
    Data collection
    Zero
    All contexts
    Input $0.14Output $0.4

    Parasail

    google/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    21 tokens/s
    Latency
    1.929 s
    Uptime
    99.7196%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.4Cached $0.06

    DeepInfra:offline

    Disabledgoogle/gemma-4-31b-it
    Context
    131,072 tokens
    Max output
    8,192 tokens
    Throughput
    15 tokens/s
    Latency
    3.969 s
    Uptime
    69.0302%
    Data collection
    Zero
    All contexts
    Input $0.27Output $0.76

    SambaNova:offline

    Disabledgoogle/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    155 tokens/s
    Latency
    1.5135 s
    Uptime
    83.6066%
    Data collection
    Zero
    All contexts
    Input $0.38Output $1.15

    Together

    google/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    19.5 tokens/s
    Latency
    0.8335 s
    Uptime
    97.8448%
    Data collection
    Zero
    All contexts
    Input $0.39Output $0.97

    ModelRun

    google/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    57 tokens/s
    Latency
    0.237 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.75Output $1Cached $0.75

    SiliconFlow:offline

    Disabledgoogle/gemma-4-31b-it
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    11 tokens/s
    Latency
    1.326 s
    Uptime
    84.009%
    Data collection
    Zero
    All contexts
    Input $0.75Output $1Cached $0.25
  72. qwen

    qwen3.6-plus

    Stable
    @qwen/qwen3.6-plus

    Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference.

    All contexts
    Input $0.325Output $1.95
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 1 for @qwen/qwen3.6-plus

    Alibaba

    qwen/qwen3.6-plus
    Context
    1,000,000 tokens
    Max output
    65,536 tokens
    Throughput
    34 tokens/s
    Latency
    0.99 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.325Output $1.95
  73. arcee-ai

    trinity-large-thinking

    Stable
    @arcee-ai/trinity-large-thinking

    Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks.

    All contexts
    Input $0.25Output $0.8Cached $0.06
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    262.1K
    ThinkingTool Calling
    Providers & technical details 1 for @arcee-ai/trinity-large-thinking

    Arcee AI

    arcee-ai/trinity-large-thinking
    Context
    262,144 tokens
    Max output
    80,000 tokens
    Throughput
    84 tokens/s
    Latency
    0.142 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.25Output $0.8Cached $0.06
  74. model-router

    claude:budget

    Stable
    @model-router/claude:budget

    Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, offering near-frontier intelligence with much lower cost and latency than larger Claude models. This model is an alias to `@anthropic/claude-4.5-haiku`.

    All contexts
    Input $1.05Output $5.25Cached $0.105
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    200K
    Image InputTool CallingModel Router
    Providers & technical details 9 for @model-router/claude:budget

    NagaAI:offline

    Disabledclaude-haiku-4.5-20251001
    Context
    200,000 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.5Output $2.5

    Azure

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    115 tokens/s
    Latency
    0.513 s
    Uptime
    100%
    Data collection
    Unknown
    All contexts
    Input $1Output $5Cached $0.1

    Amazon Bedrock

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    52 tokens/s
    Latency
    0.761 s
    Uptime
    99.9772%
    Data collection
    Zero
    All contexts
    Input $1Output $5Cached $0.1

    Google

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    68 tokens/s
    Latency
    0.795 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1Output $5Cached $0.1

    Anthropic

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    58 tokens/s
    Latency
    0.686 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1Output $5Cached $0.1

    Amazon Bedrock (US)

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Data collection
    Zero
    All contexts
    Input $1.1Output $5.5Cached $0.11

    Google (US)

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Data collection
    Zero
    All contexts
    Input $1.1Output $5.5Cached $0.11

    Amazon Bedrock (EU)

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    79 tokens/s
    Latency
    0.7215 s
    Data collection
    Zero
    All contexts
    Input $1.1Output $5.5Cached $0.11

    Google (EU)

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    78 tokens/s
    Latency
    0.3855 s
    Data collection
    Zero
    All contexts
    Input $1.1Output $5.5Cached $0.11
  75. model-router

    claude:frontier

    Stable
    @model-router/claude:frontier

    Claude Opus 5 is Anthropic's flagship model for demanding reasoning, coding, long-horizon agentic work, visual analysis, and complex professional tasks. This model is an alias to `@anthropic/claude-5-opus`.

    All contexts
    Input $5Output $25Cached $0.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 11 for @model-router/claude:frontier

    Azure (US)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    70 tokens/s
    Latency
    2.822 s
    Uptime
    100%
    Data collection
    Unknown
    All contexts
    Input $5Output $25Cached $0.5

    Claude Platform on AWS

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    38 tokens/s
    Latency
    3.96 s
    Uptime
    99.995%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Google

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    70 tokens/s
    Latency
    6.004 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Amazon Bedrock

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    70 tokens/s
    Latency
    3.948 s
    Uptime
    99.9585%
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Azure

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $5Output $25Cached $0.5

    Anthropic

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    67 tokens/s
    Latency
    3.095 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Amazon Bedrock (EU)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Google (US)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Google (EU)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    11 tokens/s
    Latency
    1.473 s
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Amazon Bedrock (US)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    56 tokens/s
    Latency
    4.68 s
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55

    Anthropic (Fast)

    anthropic/claude-opus-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    120 tokens/s
    Latency
    2.1815 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $10Output $50Cached $1
  76. model-router

    claude:frontier-mythos

    Stable
    @model-router/claude:frontier-mythos

    Claude Fable 5.1 improves on Fable 5 for agentic coding, long-running workflows, knowledge work, large refactors, visual code generation, finance, and analysis. This model is an alias to `@anthropic/claude-fable-5.1`.

    All contexts
    Input $10Output $50Cached $0.25
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 4 for @model-router/claude:frontier-mythos

    Azure

    anthropic/claude-fable-5.1
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    46 tokens/s
    Latency
    5.055 s
    Data collection
    Unknown
    All contexts
    Input $10Output $50Cached $0.25

    Anthropic

    anthropic/claude-fable-5.1
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    33 tokens/s
    Latency
    3.203 s
    Uptime
    99.8412%
    Data collection
    Prompt No Training
    All contexts
    Input $10Output $50Cached $0.25

    Amazon Bedrock

    anthropic/claude-fable-5.1
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $10Output $50Cached $0.25

    Google

    anthropic/claude-fable-5.1
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    49 tokens/s
    Latency
    4.933 s
    Uptime
    100%
    Data collection
    Unknown
    All contexts
    Input $10Output $50Cached $0.25
  77. model-router

    claude:mid

    Stable
    @model-router/claude:mid

    Claude Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. This model is an alias to `@anthropic/claude-5-sonnet`.

    All contexts
    Input $2Output $10Cached $0.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    1M
    Image InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 10 for @model-router/claude:mid

    Claude Platform on AWS

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    38 tokens/s
    Latency
    1.8165 s
    Uptime
    99.9962%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $10Cached $0.2

    Azure (US)

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    64 tokens/s
    Latency
    0.637 s
    Uptime
    100%
    Data collection
    Unknown
    All contexts
    Input $2Output $10Cached $0.2

    Azure

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    32 tokens/s
    Latency
    2.717 s
    Data collection
    Unknown
    All contexts
    Input $2Output $10Cached $0.2

    Google

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    61 tokens/s
    Latency
    2.501 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2Output $10Cached $0.2

    Amazon Bedrock

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    51 tokens/s
    Latency
    2.5175 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2Output $10Cached $0.2

    Anthropic

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    75 tokens/s
    Latency
    1.277 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $10Cached $0.2

    Amazon Bedrock (EU)

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    124 tokens/s
    Latency
    4.293 s
    Data collection
    Zero
    All contexts
    Input $2.2Output $11Cached $0.22

    Google (US)

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $2.2Output $11Cached $0.22

    Google (EU)

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    65 tokens/s
    Latency
    3.3815 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.2Output $11Cached $0.22

    Amazon Bedrock (US)

    anthropic/claude-sonnet-5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    68.5 tokens/s
    Latency
    4.339 s
    Data collection
    Zero
    All contexts
    Input $2.2Output $11Cached $0.22
  78. model-router

    deepseek:budget

    Stable
    @model-router/deepseek:budget

    DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the cost-efficient tier of the V4.1 family. This model is an alias to `@deepseek/deepseek-v4.1-flash`.

    All contexts
    Input $0.3Output $1.2Cached $0.006
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Tool CallingModel Router
    Providers & technical details 12 for @model-router/deepseek:budget

    DeepSeek

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    384,000 tokens
    Throughput
    136 tokens/s
    Latency
    1.112 s
    Uptime
    99.9981%
    Data collection
    Prompt With Training
    All contexts
    Input $0.15Output $0.6Cached $0.003

    DeepInfra:offline

    Disableddeepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    24 tokens/s
    Latency
    3.146 s
    Uptime
    91.9323%
    Data collection
    Zero
    All contexts
    Input $0.2Output $0.6Cached $0.006

    Fireworks

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    105 tokens/s
    Latency
    1.318 s
    Uptime
    99.5117%
    Data collection
    Zero
    All contexts
    Input $0.22Output $0.66Cached $0.007

    Morph

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    27 tokens/s
    Latency
    1.915 s
    Uptime
    99.473%
    Data collection
    Zero
    All contexts
    Input $0.225Output $0.9Cached $0.0225

    SiliconFlow

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    166 tokens/s
    Latency
    1.117 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Modal

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    76 tokens/s
    Latency
    1.3945 s
    Uptime
    95.4147%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.03

    Wafer

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    37 tokens/s
    Latency
    0.9055 s
    Uptime
    99.6663%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Parasail

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    92 tokens/s
    Latency
    1.186 s
    Uptime
    97.7013%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    GMICloud

    deepseek/deepseek-v4.1-flash
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    93 tokens/s
    Latency
    3.4645 s
    Uptime
    99.9774%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Io Net

    deepseek/deepseek-v4.1-flash
    Context
    262,124 tokens
    Max output
    131,072 tokens
    Throughput
    71 tokens/s
    Latency
    1.148 s
    Uptime
    99.8894%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.003

    Novita

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    130 tokens/s
    Latency
    2.059 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Venice

    deepseek/deepseek-v4.1-flash
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    97 tokens/s
    Latency
    0.979 s
    Uptime
    99.5052%
    Data collection
    Zero
    All contexts
    Input $0.375Output $1.5Cached $0.0075
  79. model-router

    deepseek:frontier

    Stable
    @model-router/deepseek:frontier

    DeepSeek V4 Pro is a large-scale Mixture-of-Experts model designed for advanced reasoning, coding, and long-horizon agent workflows with 1.6T total parameters. This model is an alias to `@deepseek/deepseek-v4-pro`.

    All contexts
    Input $1.32Output $3.96Cached $0.044
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Tool CallingModel Router
    Providers & technical details 22 for @model-router/deepseek:frontier

    NagaAI:offline

    Disableddeepseek-v4-pro
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.43Output $0.87

    Baidu

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    50 tokens/s
    Latency
    2.2955 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5782Output $1.7345Cached $0.0184

    StreamLake

    deepseek/deepseek-v4-pro-0813
    Context
    1,024,000 tokens
    Max output
    384,000 tokens
    Throughput
    62 tokens/s
    Latency
    3.1595 s
    Uptime
    99.684%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5795Output $1.7384Cached $0.0193

    Alibaba

    deepseek/deepseek-v4-pro-0813
    Context
    1,000,000 tokens
    Max output
    393,216 tokens
    Throughput
    51.5 tokens/s
    Latency
    1.23 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5808Output $1.7424Cached $0.0581

    DeepSeek

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    384,000 tokens
    Throughput
    27 tokens/s
    Latency
    1.228 s
    Uptime
    100%
    Data collection
    Prompt With Training
    All contexts
    Input $0.66Output $1.98Cached $0.022

    Ionstream

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    62 tokens/s
    Latency
    1.123 s
    Uptime
    99.9653%
    Data collection
    Zero
    All contexts
    Input $0.88Output $2.64Cached $0.088

    Novita

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    107 tokens/s
    Latency
    5.455 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.99Output $2.97Cached $0.033

    GMICloud

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    49 tokens/s
    Latency
    5.228 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.056Output $3.168Cached $0.0352

    NextBit

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    27 tokens/s
    Latency
    3.702 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.122Output $3.366Cached $0.037

    DeepInfra

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    16,384 tokens
    Throughput
    98 tokens/s
    Latency
    1.056 s
    Uptime
    99.4444%
    Data collection
    Zero
    All contexts
    Input $1.3Output $2.6Cached $0.1

    CoreWeave

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    116 tokens/s
    Latency
    0.596 s
    Uptime
    99.8523%
    Data collection
    Zero
    All contexts
    Input $1.31Output $3.96Cached $0.044

    Sail Research

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    384,000 tokens
    Throughput
    25 tokens/s
    Latency
    1.194 s
    Uptime
    98.7234%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.044

    BaseTen

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    52 tokens/s
    Latency
    0.537 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.132

    Parasail

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    54 tokens/s
    Latency
    1.1025 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.044

    Together

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    122 tokens/s
    Latency
    0.843 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.13

    DigitalOcean

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    37 tokens/s
    Latency
    0.9145 s
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.044

    SiliconFlow

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    44 tokens/s
    Latency
    1.495 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.044

    BaseTen

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    41 tokens/s
    Latency
    0.461 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.132

    Cloudflare

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    64 tokens/s
    Latency
    1.18 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.32Output $3.96Cached $0.044

    Fireworks

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    63 tokens/s
    Latency
    1.551 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.32Output $3.96Cached $0.044

    Phala

    deepseek/deepseek-v4-pro-0813
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    50 tokens/s
    Latency
    1.583 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.45Output $4.36Cached $0.15

    Venice

    deepseek/deepseek-v4-pro-0813
    Context
    1,000,000 tokens
    Max output
    32,768 tokens
    Throughput
    51 tokens/s
    Latency
    0.954 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.65Output $4.95Cached $0.165
  80. model-router

    deepseek:latest

    Stable
    @model-router/deepseek:latest

    DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the cost-efficient tier of the V4.1 family. This model is an alias to `@deepseek/deepseek-v4.1-flash`.

    All contexts
    Input $0.3Output $1.2Cached $0.006
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Tool CallingModel Router
    Providers & technical details 12 for @model-router/deepseek:latest

    DeepSeek

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    384,000 tokens
    Throughput
    136 tokens/s
    Latency
    1.112 s
    Uptime
    99.9981%
    Data collection
    Prompt With Training
    All contexts
    Input $0.15Output $0.6Cached $0.003

    DeepInfra:offline

    Disableddeepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    24 tokens/s
    Latency
    3.146 s
    Uptime
    91.9323%
    Data collection
    Zero
    All contexts
    Input $0.2Output $0.6Cached $0.006

    Fireworks

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    105 tokens/s
    Latency
    1.318 s
    Uptime
    99.5117%
    Data collection
    Zero
    All contexts
    Input $0.22Output $0.66Cached $0.007

    Morph

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    27 tokens/s
    Latency
    1.915 s
    Uptime
    99.473%
    Data collection
    Zero
    All contexts
    Input $0.225Output $0.9Cached $0.0225

    SiliconFlow

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    166 tokens/s
    Latency
    1.117 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Modal

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    76 tokens/s
    Latency
    1.3945 s
    Uptime
    95.4147%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.03

    Wafer

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    37 tokens/s
    Latency
    0.9055 s
    Uptime
    99.6663%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Parasail

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    92 tokens/s
    Latency
    1.186 s
    Uptime
    97.7013%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    GMICloud

    deepseek/deepseek-v4.1-flash
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    93 tokens/s
    Latency
    3.4645 s
    Uptime
    99.9774%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Io Net

    deepseek/deepseek-v4.1-flash
    Context
    262,124 tokens
    Max output
    131,072 tokens
    Throughput
    71 tokens/s
    Latency
    1.148 s
    Uptime
    99.8894%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.003

    Novita

    deepseek/deepseek-v4.1-flash
    Context
    1,048,576 tokens
    Max output
    393,216 tokens
    Throughput
    130 tokens/s
    Latency
    2.059 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.006

    Venice

    deepseek/deepseek-v4.1-flash
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    97 tokens/s
    Latency
    0.979 s
    Uptime
    99.5052%
    Data collection
    Zero
    All contexts
    Input $0.375Output $1.5Cached $0.0075
  81. model-router

    glm:budget

    Stable
    @model-router/glm:budget

    GLM-5.3-Flash is Z.ai's efficient native multimodal model for coding and long-horizon agent tasks, with image and video understanding and a 1M-token context window. This model is an alias to `@z-ai/glm-5.3-flash`.

    All contexts
    Input $0.15Output $0.5Cached $0.03
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1.3M
    Image InputVideo InputThinkingTool CallingModel Router
    Providers & technical details 26 for @model-router/glm:budget

    DeepInfra

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    18 tokens/s
    Latency
    1.9565 s
    Uptime
    99.1382%
    Data collection
    Zero
    All contexts
    Input $0.075Output $0.25Cached $0.015

    Relace

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    35 tokens/s
    Latency
    1.26 s
    Uptime
    99.9482%
    Data collection
    Zero
    All contexts
    Input $0.09Output $0.3Cached $0.018

    Morph

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.35Cached $0.02

    Wafer

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    9 tokens/s
    Latency
    0.871 s
    Uptime
    99.9709%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.35Cached $0.02

    StreamLake

    z-ai/glm-5.3-flash
    Context
    1,024,000 tokens
    Max output
    128,000 tokens
    Throughput
    43 tokens/s
    Latency
    1.529 s
    Uptime
    99.447%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1124Output $0.3745Cached $0.0225

    GMICloud

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    26 tokens/s
    Latency
    2.083 s
    Uptime
    99.5128%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1125Output $0.375Cached $0.0225

    Reka

    z-ai/glm-5.3-flash
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    14 tokens/s
    Latency
    1.608 s
    Uptime
    99.8756%
    Data collection
    Zero
    All contexts
    Input $0.132Output $0.44Cached $0.0264

    Novita

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    32 tokens/s
    Latency
    3.292 s
    Uptime
    99.7808%
    Data collection
    Zero
    All contexts
    Input $0.132Output $0.44Cached $0.0264

    Makora

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    69 tokens/s
    Latency
    0.831 s
    Uptime
    99.4227%
    Data collection
    Zero
    All contexts
    Input $0.14Output $0.47Cached $0.024

    Crusoe:offline

    Disabledz-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    86 tokens/s
    Latency
    0.876 s
    Uptime
    93.681%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    CoreWeave

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    31 tokens/s
    Latency
    1.239 s
    Uptime
    99.3789%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.05

    Sail Research

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    24 tokens/s
    Latency
    1.899 s
    Uptime
    99.2481%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Fireworks

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    78 tokens/s
    Latency
    0.953 s
    Uptime
    99.8622%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Phala

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    33 tokens/s
    Latency
    3.249 s
    Uptime
    99.7407%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Friendli

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    70 tokens/s
    Latency
    0.7515 s
    Uptime
    95.8315%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $0.5Cached $0.03

    SiliconFlow

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    28 tokens/s
    Latency
    1.656 s
    Uptime
    99.8555%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    DigitalOcean

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    17 tokens/s
    Latency
    2.613 s
    Uptime
    97.3703%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Together

    z-ai/glm-5.3-flash
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    33.5 tokens/s
    Latency
    0.572 s
    Uptime
    99.906%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Parasail

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    60 tokens/s
    Latency
    1.219 s
    Uptime
    99.6513%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    BaseTen

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    135 tokens/s
    Latency
    0.7615 s
    Uptime
    99.9464%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Venice:offline

    Disabledz-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    20 tokens/s
    Latency
    2.583 s
    Uptime
    93.5544%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Io Net

    z-ai/glm-5.3-flash
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    30 tokens/s
    Latency
    1.24 s
    Uptime
    99.8962%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Cloudflare

    z-ai/glm-5.3-flash
    Context
    1,310,720 tokens
    Max output
    1,179,648 tokens
    Throughput
    49 tokens/s
    Latency
    0.8855 s
    Uptime
    99.9447%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $0.5Cached $0.03

    Z.AI

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    53 tokens/s
    Latency
    2.516 s
    Uptime
    99.1417%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.5Cached $0.03

    NextBit

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    128,000 tokens
    Throughput
    41 tokens/s
    Latency
    2.4155 s
    Uptime
    99.8763%
    Data collection
    Zero
    All contexts
    Input $0.177Output $0.59Cached $0.036

    Modal

    z-ai/glm-5.3-flash
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    74 tokens/s
    Latency
    0.6995 s
    Uptime
    99.797%
    Data collection
    Zero
    All contexts
    Input $0.45Output $1.5Cached $0.09
  82. model-router

    glm:frontier

    Stable
    @model-router/glm:frontier

    GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window. This model is an alias to `@z-ai/glm-5.3`.

    All contexts
    Input $1.4Output $4.4Cached $0.26
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1.3M
    ThinkingTool CallingModel Router
    Providers & technical details 26 for @model-router/glm:frontier

    Inceptron

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    38 tokens/s
    Latency
    0.611 s
    Uptime
    99.2912%
    Data collection
    Zero
    All contexts
    Input $0.8727Output $3.36Cached $0.1639

    Reka

    z-ai/glm-5.3
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    58 tokens/s
    Latency
    1.507 s
    Uptime
    99.5154%
    Data collection
    Zero
    All contexts
    Input $0.936Output $3.168Cached $0.1872

    DigitalOcean

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    43 tokens/s
    Latency
    9.611 s
    Uptime
    97.4692%
    Data collection
    Zero
    All contexts
    Input $0.95Output $3.4Cached $0.2

    Morph

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    79 tokens/s
    Latency
    1.014 s
    Uptime
    98.7516%
    Data collection
    Zero
    All contexts
    Input $1Output $3.41Cached $0.2

    Phala

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    57 tokens/s
    Latency
    1.954 s
    Uptime
    98.5335%
    Data collection
    Zero
    All contexts
    Input $1.05Output $3.3Cached $0.195

    Novita

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    42 tokens/s
    Latency
    2.881 s
    Uptime
    99.929%
    Data collection
    Zero
    All contexts
    Input $1.092Output $3.432Cached $0.2028

    GMICloud

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    48 tokens/s
    Latency
    1.196 s
    Uptime
    99.2525%
    Data collection
    Prompt No Training
    All contexts
    Input $1.12Output $3.52Cached $0.208

    AkashML

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    53.5 tokens/s
    Latency
    4.0715 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.17Output $3.96Cached $0.234

    Decart

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    94 tokens/s
    Latency
    1.072 s
    Uptime
    99.9573%
    Data collection
    Zero
    All contexts
    Input $1.19Output $3.74Cached $0.1955

    Wafer

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    86 tokens/s
    Latency
    2.81 s
    Uptime
    98.6798%
    Data collection
    Zero
    All contexts
    Input $1.19Output $4.4Cached $0.26

    DeepInfra

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    28 tokens/s
    Latency
    9.307 s
    Uptime
    98.4858%
    Data collection
    Zero
    All contexts
    Input $1.2Output $4Cached $0.12

    Sail Research

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    66 tokens/s
    Latency
    1.8615 s
    Uptime
    99.7837%
    Data collection
    Zero
    All contexts
    Input $1.2572Output $3.9512Cached $0.2335

    Friendli

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    125 tokens/s
    Latency
    1.559 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.26Output $3.96Cached $0.234

    Makora

    z-ai/glm-5.3
    Context
    980,000 tokens
    Max output
    128,000 tokens
    Throughput
    111 tokens/s
    Latency
    1.423 s
    Data collection
    Zero
    All contexts
    Input $1.35Output $4.4Cached $0.23

    Crusoe

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    152 tokens/s
    Latency
    0.545 s
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Venice

    z-ai/glm-5.3
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    39.5 tokens/s
    Latency
    4.3025 s
    Uptime
    97.7465%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    SiliconFlow

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    36 tokens/s
    Latency
    1.2375 s
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Together

    z-ai/glm-5.3
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    123 tokens/s
    Latency
    0.591 s
    Uptime
    97.4216%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Parasail

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    60 tokens/s
    Latency
    1.273 s
    Uptime
    99.9025%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Modal

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    88 tokens/s
    Latency
    3.059 s
    Uptime
    99.8686%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    BaseTen:offline

    Disabledz-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    84 tokens/s
    Latency
    1.531 s
    Uptime
    91.8782%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.14

    Fireworks

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    54 tokens/s
    Latency
    1.691 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Cloudflare

    z-ai/glm-5.3
    Context
    1,310,720 tokens
    Max output
    1,179,648 tokens
    Throughput
    50 tokens/s
    Latency
    5.05 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    AtlasCloud

    z-ai/glm-5.3
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    74 tokens/s
    Latency
    4.565 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Z.AI

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    57 tokens/s
    Latency
    2.587 s
    Uptime
    99.6917%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    BaseTen

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    60 tokens/s
    Latency
    0.927 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.1Output $6.6Cached $0.21
  83. model-router

    glm:latest

    Stable
    @model-router/glm:latest

    GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window. This model is an alias to `@z-ai/glm-5.3`.

    All contexts
    Input $1.4Output $4.4Cached $0.26
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1.3M
    ThinkingTool CallingModel Router
    Providers & technical details 26 for @model-router/glm:latest

    Inceptron

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    38 tokens/s
    Latency
    0.611 s
    Uptime
    99.2912%
    Data collection
    Zero
    All contexts
    Input $0.8727Output $3.36Cached $0.1639

    Reka

    z-ai/glm-5.3
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    58 tokens/s
    Latency
    1.507 s
    Uptime
    99.5154%
    Data collection
    Zero
    All contexts
    Input $0.936Output $3.168Cached $0.1872

    DigitalOcean

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    43 tokens/s
    Latency
    9.611 s
    Uptime
    97.4692%
    Data collection
    Zero
    All contexts
    Input $0.95Output $3.4Cached $0.2

    Morph

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    79 tokens/s
    Latency
    1.014 s
    Uptime
    98.7516%
    Data collection
    Zero
    All contexts
    Input $1Output $3.41Cached $0.2

    Phala

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    57 tokens/s
    Latency
    1.954 s
    Uptime
    98.5335%
    Data collection
    Zero
    All contexts
    Input $1.05Output $3.3Cached $0.195

    Novita

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    42 tokens/s
    Latency
    2.881 s
    Uptime
    99.929%
    Data collection
    Zero
    All contexts
    Input $1.092Output $3.432Cached $0.2028

    GMICloud

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    48 tokens/s
    Latency
    1.196 s
    Uptime
    99.2525%
    Data collection
    Prompt No Training
    All contexts
    Input $1.12Output $3.52Cached $0.208

    AkashML

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    53.5 tokens/s
    Latency
    4.0715 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.17Output $3.96Cached $0.234

    Decart

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    94 tokens/s
    Latency
    1.072 s
    Uptime
    99.9573%
    Data collection
    Zero
    All contexts
    Input $1.19Output $3.74Cached $0.1955

    Wafer

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    86 tokens/s
    Latency
    2.81 s
    Uptime
    98.6798%
    Data collection
    Zero
    All contexts
    Input $1.19Output $4.4Cached $0.26

    DeepInfra

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    28 tokens/s
    Latency
    9.307 s
    Uptime
    98.4858%
    Data collection
    Zero
    All contexts
    Input $1.2Output $4Cached $0.12

    Sail Research

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    66 tokens/s
    Latency
    1.8615 s
    Uptime
    99.7837%
    Data collection
    Zero
    All contexts
    Input $1.2572Output $3.9512Cached $0.2335

    Friendli

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    125 tokens/s
    Latency
    1.559 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.26Output $3.96Cached $0.234

    Makora

    z-ai/glm-5.3
    Context
    980,000 tokens
    Max output
    128,000 tokens
    Throughput
    111 tokens/s
    Latency
    1.423 s
    Data collection
    Zero
    All contexts
    Input $1.35Output $4.4Cached $0.23

    Crusoe

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    152 tokens/s
    Latency
    0.545 s
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Venice

    z-ai/glm-5.3
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    39.5 tokens/s
    Latency
    4.3025 s
    Uptime
    97.7465%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    SiliconFlow

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    36 tokens/s
    Latency
    1.2375 s
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Together

    z-ai/glm-5.3
    Context
    1,048,575 tokens
    Max output
    943,717 tokens
    Throughput
    123 tokens/s
    Latency
    0.591 s
    Uptime
    97.4216%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Parasail

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    60 tokens/s
    Latency
    1.273 s
    Uptime
    99.9025%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Modal

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    88 tokens/s
    Latency
    3.059 s
    Uptime
    99.8686%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    BaseTen:offline

    Disabledz-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    84 tokens/s
    Latency
    1.531 s
    Uptime
    91.8782%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.14

    Fireworks

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    54 tokens/s
    Latency
    1.691 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Cloudflare

    z-ai/glm-5.3
    Context
    1,310,720 tokens
    Max output
    1,179,648 tokens
    Throughput
    50 tokens/s
    Latency
    5.05 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    AtlasCloud

    z-ai/glm-5.3
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    74 tokens/s
    Latency
    4.565 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.4Output $4.4Cached $0.26

    Z.AI

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    57 tokens/s
    Latency
    2.587 s
    Uptime
    99.6917%
    Data collection
    Zero
    All contexts
    Input $1.4Output $4.4Cached $0.26

    BaseTen

    z-ai/glm-5.3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    60 tokens/s
    Latency
    0.927 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.1Output $6.6Cached $0.21
  84. model-router

    google:budget

    Stable
    @model-router/google:budget

    Gemini 3.5 Flash-Lite is Google's most cost-efficient general availability model, optimized for high-volume agentic tasks, translation, and simple data processing. This model is an alias to `@google/gemini-3.5-flash-lite`.

    All contexts
    Input $0.315Output $2.625Cached $0.0315Audio input $0.315
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    Medium
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 8 for @model-router/google:budget

    Google (Flex)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    109 tokens/s
    Latency
    6.608 s
    Uptime
    96.6781%
    Data collection
    Zero
    All contexts
    Input $0.15Output $1.25Cached $0.015Audio input $0.15

    Google AI Studio (Flex)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    167.5 tokens/s
    Latency
    1.9205 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $1.25Cached $0.015Audio input $0.15

    Google AI Studio

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    77 tokens/s
    Latency
    0.512 s
    Uptime
    99.9918%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $2.5Cached $0.03Audio input $0.3

    Google

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    70 tokens/s
    Latency
    0.579 s
    Uptime
    99.9135%
    Data collection
    Zero
    All contexts
    Input $0.3Output $2.5Cached $0.03Audio input $0.3

    Google (EU)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Zero
    All contexts
    Input $0.33Output $2.75Cached $0.033Audio input $0.33

    Google (US)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Zero
    All contexts
    Input $0.33Output $2.75Cached $0.033Audio input $0.33

    Google (Priority)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    10 tokens/s
    Latency
    0.923 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.54Output $4.5Cached $0.054Audio input $0.54

    Google AI Studio (Priority)

    google/gemini-3.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    10 tokens/s
    Latency
    0.932 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.54Output $4.5Cached $0.054Audio input $0.54
  85. model-router

    google:frontier

    PreviewStable
    @model-router/google:frontier

    Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. This model is an alias to `@google/gemini-3.1-pro`.

    All contexts
    Input $2Output $12Cached $0.2Audio input $2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 7 for @model-router/google:frontier

    NagaAI:offline

    Disabledgemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    All contexts
    Input $1Output $6

    Google (Flex):offline

    Disabledgoogle/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Zero
    All contexts
    Input $1Output $6Cached $0.1Audio input $1

    Google AI Studio (Flex)

    google/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    62 tokens/s
    Latency
    2.767 s
    Data collection
    Prompt No Training
    All contexts
    Input $1Output $6Cached $0.1Audio input $1

    Google

    google/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    83 tokens/s
    Latency
    2.7025 s
    Uptime
    97.3863%
    Data collection
    Zero
    All contexts
    Input $2Output $12Cached $0.2Audio input $2

    Google AI Studio

    google/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    96 tokens/s
    Latency
    3.135 s
    Uptime
    99.7345%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $12Cached $0.2Audio input $2

    Google (Priority)

    google/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Zero
    All contexts
    Input $3.6Output $21.6Cached $0.36Audio input $3.6

    Google AI Studio (Priority)

    google/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $3.6Output $21.6Cached $0.36Audio input $3.6
  86. model-router

    google:mid

    Stable
    @model-router/google:mid

    Gemini 3.8 Flash is Google's most intelligent Flash model, with significant gains over 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. This model is an alias to `@google/gemini-3.8-flash`.

    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 6 for @model-router/google:mid

    Google AI Studio (Flex)

    google/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    174 tokens/s
    Latency
    1.3775 s
    Uptime
    99.8152%
    Data collection
    Prompt No Training
    All contexts
    Input $0.375Output $1.875Cached $0.0375Audio input $0.375

    Google (Flex):offline

    Disabledgoogle/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    37 tokens/s
    Latency
    12.23 s
    Uptime
    94.7605%
    Data collection
    Zero
    All contexts
    Input $0.375Output $1.875Cached $0.0375Audio input $0.375

    Google AI Studio

    google/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    187 tokens/s
    Latency
    1.393 s
    Uptime
    99.8256%
    Data collection
    Prompt No Training
    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75

    Google

    google/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    78 tokens/s
    Latency
    2.2695 s
    Uptime
    98.5903%
    Data collection
    Zero
    All contexts
    Input $0.75Output $3.75Cached $0.075Audio input $0.75

    Google AI Studio (Priority)

    google/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    28 tokens/s
    Latency
    1.7825 s
    Uptime
    98.4197%
    Data collection
    Prompt No Training
    All contexts
    Input $1.35Output $6.75Cached $0.135Audio input $1.35

    Google (Priority)

    google/gemini-3.8-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    20 tokens/s
    Latency
    2.0645 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.35Output $6.75Cached $0.135Audio input $1.35
  87. model-router

    grok:latest

    Stable
    @model-router/grok:latest

    Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM. This model is an alias to `@x-ai/grok-4.6`.

    All contexts
    Input $2Output $6Cached $0.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    500K
    Image InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 5 for @model-router/grok:latest

    xAI (ZDR)

    x-ai/grok-4.6
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    53 tokens/s
    Latency
    0.673 s
    Uptime
    99.7449%
    Data collection
    Zero
    All contexts
    Input $2Output $6Cached $0.5

    xAI

    x-ai/grok-4.6
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    55 tokens/s
    Latency
    1.247 s
    Uptime
    99.8829%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $6Cached $0.5

    Amazon Bedrock (US)

    x-ai/grok-4.6
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Data collection
    Zero
    All contexts
    Input $2.2Output $6.6Cached $0.55

    xAI (Priority) (ZDR)

    x-ai/grok-4.6
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    41 tokens/s
    Latency
    1.84 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $4Output $12Cached $1

    xAI (Priority)

    x-ai/grok-4.6
    Context
    500,000 tokens
    Max output
    450,000 tokens
    Throughput
    60 tokens/s
    Latency
    1.228 s
    Data collection
    Prompt No Training
    All contexts
    Input $4Output $12Cached $1
  88. model-router

    kimi:latest

    Stable
    @model-router/kimi:latest

    Kimi K3 is Moonshot AI's ultra-large-scale, open-weight multimodal reasoning model for complex coding, knowledge work, and long-horizon agentic workflows. This model is an alias to `@moonshotai/kimi-k3`.

    All contexts
    Input $3Output $15Cached $0.3
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool CallingModel Router
    Providers & technical details 20 for @model-router/kimi:latest

    NagaAI:offline

    Disabledkimi-k3
    Context
    1,048,576 tokens
    Max output
    1,048,576 tokens
    All contexts
    Input $1.5Output $7.5

    Morph

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    11 tokens/s
    Latency
    3.127 s
    Uptime
    99.4364%
    Data collection
    Zero
    All contexts
    Input $2.375Output $13.3Cached $0.2755

    Relace

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    45 tokens/s
    Latency
    1.6025 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.4Output $12Cached $0.24

    Makora:offline

    Disabledmoonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    33 tokens/s
    Latency
    1.071 s
    Uptime
    84.5455%
    Data collection
    Zero
    All contexts
    Input $2.55Output $12.75Cached $0.256

    DigitalOcean

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    20 tokens/s
    Latency
    2.7115 s
    Uptime
    99.5723%
    Data collection
    Zero
    All contexts
    Input $2.55Output $12.95Cached $0.285

    Sail Research

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    71 tokens/s
    Latency
    1.465 s
    Uptime
    99.9739%
    Data collection
    Zero
    All contexts
    Input $2.6481Output $13.2827Cached $0.3026

    Phala

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    33 tokens/s
    Latency
    2.5875 s
    Uptime
    96.3516%
    Data collection
    Zero
    All contexts
    Input $2.85Output $14.25Cached $0.285

    DeepInfra

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    16,384 tokens
    Throughput
    17.5 tokens/s
    Latency
    4.726 s
    Uptime
    99.8598%
    Data collection
    Zero
    All contexts
    Input $2.85Output $14.25Cached $0.285

    Wafer

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    30 tokens/s
    Latency
    0.7305 s
    Uptime
    99.4998%
    Data collection
    Zero
    All contexts
    Input $3Output $12.75Cached $0.3

    Chutes

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    24 tokens/s
    Latency
    2.482 s
    Data collection
    Prompt No Training
    All contexts
    Input $3Output $15Cached $0.3

    Parasail

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    30 tokens/s
    Latency
    1.007 s
    Uptime
    99.7284%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Modal

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    75.5 tokens/s
    Latency
    1.3635 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Together

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    35 tokens/s
    Latency
    1.434 s
    Uptime
    99.7069%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Fireworks

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    42 tokens/s
    Latency
    1.7035 s
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    BaseTen

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    262,144 tokens
    Throughput
    69 tokens/s
    Latency
    1.696 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Moonshot AI

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    24 tokens/s
    Latency
    5.042 s
    Uptime
    99.0595%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Fireworks (US):offline

    Disabledmoonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    33 tokens/s
    Latency
    3.0655 s
    Uptime
    87.6529%
    Data collection
    Zero
    All contexts
    Input $3.3Output $16.5Cached $0.33

    Alibaba

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    35 tokens/s
    Latency
    1.727 s
    Uptime
    99.1111%
    Data collection
    Prompt No Training
    All contexts
    Input $3.45Output $17.25Cached $0.345

    Fireworks (Fast)

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    88 tokens/s
    Latency
    0.6455 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $4.5Output $22.5Cached $0.45

    Morph (Fast)

    moonshotai/kimi-k3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    12 tokens/s
    Latency
    2.125 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $6Output $22.5Cached $0.6
  89. model-router

    mercury:latest

    Stable
    @model-router/mercury:latest

    Mercury 2.5 is Inception's fastest diffusion reasoning LLM, producing and refining multiple tokens in parallel for agentic and coding workloads. This model is an alias to `@inception/mercury-2.5`.

    All contexts
    Input $0.04Output $0.15Cached $0.004
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    260K
    ThinkingTool CallingModel RouterDiffusion
    Providers & technical details 1 for @model-router/mercury:latest

    Inception

    inception/mercury-2.5
    Context
    260,000 tokens
    Max output
    65,536 tokens
    Throughput
    76 tokens/s
    Latency
    0.677 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.04Output $0.15Cached $0.004
  90. model-router

    meta:latest

    Stable
    @model-router/meta:latest

    Muse Spark 1.3 is Meta's multimodal reasoning model for long-running agentic, multi-agent, and coding workflows, with emphasis on tracking information across extended tasks and concise execution. This model is an alias to `@meta/muse-spark-1.3`.

    All contexts
    Input $1.25Output $4.25Cached $0.15
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 2 for @model-router/meta:latest

    Meta (Alt route)

    muse-spark-1.3
    Context
    1,000,000 tokens
    Max output
    1,000,000 tokens
    Data collection
    Prompt With Training
    All contexts
    Input $1.25Output $4.25Cached $0.15

    Meta

    meta/muse-spark-1.3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    105 tokens/s
    Latency
    2.2415 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $4.25Cached $0.15
  91. model-router

    minimax:latest

    Stable
    @model-router/minimax:latest

    MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. This model is an alias to `@minimax/m3`.

    All contexts
    Input $0.3Output $1.2Cached $0.06
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Image InputVideo InputThinkingTool CallingModel Router
    Providers & technical details 13 for @model-router/minimax:latest

    NagaAI:offline

    Disabledminimax-m3
    Context
    524,288 tokens
    Max output
    524,288 tokens
    All contexts
    Input $0.15Output $0.6

    CoreWeave

    minimax/minimax-m3
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    46 tokens/s
    Latency
    0.576 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.23Output $0.96Cached $0.05

    GMICloud

    minimax/minimax-m3
    Context
    1,048,576 tokens
    Max output
    524,288 tokens
    Throughput
    57 tokens/s
    Latency
    1.066 s
    Uptime
    99.6851%
    Data collection
    Prompt No Training
    All contexts
    Input $0.24Output $0.96Cached $0.048

    DeepInfra

    minimax/minimax-m3
    Context
    524,288 tokens
    Max output
    512,000 tokens
    Throughput
    23.5 tokens/s
    Latency
    2.3585 s
    Uptime
    98.9899%
    Data collection
    Zero
    All contexts
    Input $0.28Output $1.1Cached $0.056

    StreamLake

    minimax/minimax-m3
    Context
    1,000,000 tokens
    Max output
    512,000 tokens
    Throughput
    63 tokens/s
    Latency
    1.333 s
    Uptime
    99.2432%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Venice:offline

    Disabledminimax/minimax-m3
    Context
    524,288 tokens
    Max output
    65,536 tokens
    Throughput
    67 tokens/s
    Latency
    0.89 s
    Uptime
    85.879%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Together

    minimax/minimax-m3
    Context
    524,288 tokens
    Max output
    471,859 tokens
    Throughput
    45 tokens/s
    Latency
    1.416 s
    Uptime
    99.6058%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Parasail

    minimax/minimax-m3
    Context
    1,048,576 tokens
    Max output
    524,288 tokens
    Throughput
    75 tokens/s
    Latency
    0.4835 s
    Uptime
    99.9324%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.06

    AtlasCloud

    minimax/minimax-m3
    Context
    524,300 tokens
    Max output
    524,288 tokens
    Throughput
    110 tokens/s
    Latency
    2.709 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Novita

    minimax/minimax-m3
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    73 tokens/s
    Latency
    1.595 s
    Uptime
    99.8412%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Minimax

    minimax/minimax-m3
    Context
    524,288 tokens
    Max output
    512,000 tokens
    Throughput
    96 tokens/s
    Latency
    0.817 s
    Uptime
    99.1853%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.06

    SambaNova

    minimax/minimax-m3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    136 tokens/s
    Latency
    1.819 s
    Uptime
    99.3007%
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.4

    ModelRun

    minimax/minimax-m3
    Context
    1,048,576 tokens
    Max output
    943,718 tokens
    Throughput
    126 tokens/s
    Latency
    0.658 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.75Output $3Cached $0.15
  92. model-router

    openai:budget

    Stable
    @model-router/openai:budget

    GPT-5.6 Luna is OpenAI's fast, cost-efficient GPT-5.6 model for high-volume chat, classification, and lightweight agentic workflows. This model is an alias to `@openai/gpt-5.6-luna`.

    All contexts
    Input $0.22Output $1.32Cached $0.022
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    1.1M
    Image InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 7 for @model-router/openai:budget

    OpenAI (Flex)

    openai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    66 tokens/s
    Latency
    3.831 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1Output $0.6Cached $0.01

    Azure:offline

    Disabledopenai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    68 tokens/s
    Latency
    3.562 s
    Uptime
    85.2456%
    Data collection
    Zero
    All contexts
    Input $0.2Output $1.2Cached $0.02

    OpenAI

    openai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    68 tokens/s
    Latency
    3.372 s
    Uptime
    99.8942%
    Data collection
    Prompt No Training
    All contexts
    Input $0.2Output $1.2Cached $0.02

    Azure (US):offline

    Disabledopenai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    62 tokens/s
    Latency
    4.0455 s
    Uptime
    80.9646%
    Data collection
    Zero
    All contexts
    Input $0.22Output $1.32Cached $0.022

    Amazon Bedrock (US)

    openai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    124 tokens/s
    Latency
    0.543 s
    Uptime
    100%
    Data collection
    Unknown
    All contexts
    Input $0.22Output $1.32Cached $0.022

    Azure (EU)

    openai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    61 tokens/s
    Latency
    1.3555 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.22Output $1.32Cached $0.022

    OpenAI (Fast)

    openai/gpt-5.6-luna
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    87 tokens/s
    Latency
    1.662 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.4Output $2.4Cached $0.04
  93. model-router

    openai:frontier

    Stable
    @model-router/openai:frontier

    GPT-6 Astra is a frontier reasoning model from OpenAI, suited for complex reasoning, coding, and agentic workflows. This model is an alias to `@openai/gpt-6-astra`.

    All contexts
    Input $10Output $50Cached $1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Highest
    Context up to
    1.1M
    Image InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 5 for @model-router/openai:frontier

    OpenAI (Flex)

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    58 tokens/s
    Latency
    3.2165 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Azure

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    23 tokens/s
    Latency
    10.211 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $10Output $50Cached $1

    OpenAI

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    33 tokens/s
    Latency
    2.641 s
    Uptime
    99.5477%
    Data collection
    Prompt No Training
    All contexts
    Input $10Output $50Cached $1

    Azure (US)

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    23 tokens/s
    Latency
    4.8155 s
    Data collection
    Zero
    All contexts
    Input $11Output $55Cached $1.1

    OpenAI (Fast)

    openai/gpt-6-astra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    57 tokens/s
    Latency
    3.992 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $20Output $100Cached $2
  94. model-router

    openai:mid

    Stable
    @model-router/openai:mid

    GPT-5.6 Terra is OpenAI's balanced GPT-5.6 model for everyday coding, reasoning, and agentic tasks. This model is an alias to `@openai/gpt-5.6-terra`.

    All contexts
    Input $2.2Output $13.2Cached $0.22
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1.1M
    Image InputThinkingTool CallingModel RouterFile Input
    Providers & technical details 7 for @model-router/openai:mid

    OpenAI (Flex)

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    82 tokens/s
    Latency
    2.1045 s
    Data collection
    Prompt No Training
    All contexts
    Input $1Output $6Cached $0.1

    Azure

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    54 tokens/s
    Latency
    3.578 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2Output $12Cached $0.2

    OpenAI

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    59 tokens/s
    Latency
    1.9355 s
    Uptime
    98.6473%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $12Cached $0.2

    Azure (US)

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    104 tokens/s
    Latency
    13.961 s
    Data collection
    Zero
    All contexts
    Input $2.2Output $13.2Cached $0.22

    Amazon Bedrock (US)

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $2.2Output $13.2Cached $0.22

    Azure (EU)

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    30 tokens/s
    Latency
    1.581 s
    Data collection
    Zero
    All contexts
    Input $2.2Output $13.2Cached $0.22

    OpenAI (Fast)

    openai/gpt-5.6-terra
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    36 tokens/s
    Latency
    1.2195 s
    Data collection
    Prompt No Training
    All contexts
    Input $4Output $24Cached $0.4
  95. model-router

    qwen:budget

    Stable
    @model-router/qwen:budget

    Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis. This model is an alias to `@qwen/qwen3.8-flash`.

    All contexts
    Input $0.15Output $0.47Cached $0.016
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    1M
    Image InputVideo InputThinkingTool CallingModel Router
    Providers & technical details 2 for @model-router/qwen:budget

    Makora

    qwen/qwen3.8-flash
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    92 tokens/s
    Latency
    0.942 s
    Uptime
    99.7969%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.47Cached $0.016

    Alibaba

    qwen/qwen3.8-flash
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    51 tokens/s
    Latency
    1.556 s
    Uptime
    99.5586%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $0.47Cached $0.016
  96. model-router

    qwen:frontier

    Stable
    @model-router/qwen:frontier

    Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding, coding, and agentic workflows. This model is an alias to `@qwen/qwen3.8-max`.

    All contexts
    Input $2Output $6Cached $0.25
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1M
    Image InputVideo InputThinkingTool CallingModel Router
    Providers & technical details 1 for @model-router/qwen:frontier

    Alibaba

    qwen/qwen3.8-max
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    37 tokens/s
    Latency
    1.769 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $6Cached $0.25
  97. model-router

    qwen:mid

    Stable
    @model-router/qwen:mid

    Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interactive agent tasks, with flexible thinking that can be enabled or disabled. This model is an alias to `@qwen/qwen3.8-27b`.

    All contexts
    Input $0.3333Output $2.775Cached $0.05
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    1M
    Image InputVideo InputThinkingTool CallingModel Router
    Providers & technical details 15 for @model-router/qwen:mid

    Darkbloom

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    32,768 tokens
    Throughput
    26 tokens/s
    Latency
    1.9755 s
    Uptime
    99.7451%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $2

    DekaLLM

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    63 tokens/s
    Latency
    0.679 s
    Uptime
    97.5967%
    Data collection
    Prompt No Training
    All contexts
    Input $0.2Output $2.5Cached $0.05

    Reka

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    53 tokens/s
    Latency
    0.576 s
    Uptime
    99.9229%
    Data collection
    Zero
    All contexts
    Input $0.214Output $2.55Cached $0.15

    Parasail

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    56 tokens/s
    Latency
    0.801 s
    Uptime
    99.965%
    Data collection
    Zero
    All contexts
    Input $0.24Output $2.2Cached $0.05

    AkashML

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    52 tokens/s
    Latency
    0.52 s
    Uptime
    99.954%
    Data collection
    Zero
    All contexts
    Input $0.25Output $2.2Cached $0.05

    Mancer 2

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    40 tokens/s
    Latency
    0.899 s
    Uptime
    99.7106%
    Data collection
    Zero
    All contexts
    Input $0.25Output $2.75

    Io Net

    qwen/qwen3.8-27b
    Context
    65,536 tokens
    Max output
    58,982 tokens
    Throughput
    24 tokens/s
    Latency
    0.692 s
    Uptime
    98.2318%
    Data collection
    Zero
    All contexts
    Input $0.3Output $2.8Cached $0.18

    Phala

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    46 tokens/s
    Latency
    0.4635 s
    Uptime
    99.2896%
    Data collection
    Zero
    All contexts
    Input $0.3Output $3Cached $0.05

    Chutes

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    31 tokens/s
    Latency
    2.122 s
    Uptime
    99.8926%
    Data collection
    Prompt No Training
    All contexts
    Input $0.32Output $2.5Cached $0.032

    Ionstream

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    47 tokens/s
    Latency
    0.628 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.35Output $2.55Cached $0.05

    CoreWeave

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    61 tokens/s
    Latency
    0.506 s
    Uptime
    99.9358%
    Data collection
    Zero
    All contexts
    Input $0.4Output $3Cached $0.15

    Novita

    qwen/qwen3.8-27b
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    46 tokens/s
    Latency
    2.982 s
    Uptime
    99.7294%
    Data collection
    Zero
    All contexts
    Input $0.42Output $3Cached $0.085

    Alibaba

    qwen/qwen3.8-27b
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    Throughput
    45 tokens/s
    Latency
    0.899 s
    Uptime
    99.8596%
    Data collection
    Prompt No Training
    All contexts
    Input $0.425Output $2.55Cached $0.085

    Cloudflare

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    36 tokens/s
    Latency
    0.5845 s
    Uptime
    98.419%
    Data collection
    Prompt No Training
    All contexts
    Input $0.45Output $3.2Cached $0.05

    Venice

    qwen/qwen3.8-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    70 tokens/s
    Latency
    1.02 s
    Uptime
    99.3289%
    Data collection
    Zero
    All contexts
    Input $0.45Output $3.2
  98. model-router

    xiaomi:budget

    Stable
    @model-router/xiaomi:budget

    MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. This model is an alias to `@xiaomi/mimo-v2.5`.

    All contexts
    Input $0.168Output $0.336Cached $0.003
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingModel Router
    Providers & technical details 6 for @model-router/xiaomi:budget

    Xiaomi:offline

    Disabledmimo-v2.5
    Context
    1,050,000 tokens
    Max output
    131,072 tokens

    Pricing not published.

    GMICloud:offline

    Disabledxiaomi/mimo-v2.5
    Context
    1,050,000 tokens
    Max output
    945,000 tokens
    Throughput
    2 tokens/s
    Latency
    4.4005 s
    Uptime
    66.619%
    Data collection
    Prompt No Training
    All contexts
    Input $0.119Output $0.238Cached $0.0026

    DeepInfra

    xiaomi/mimo-v2.5
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    13 tokens/s
    Latency
    1.766 s
    Uptime
    99.8196%
    Data collection
    Zero
    All contexts
    Input $0.133Output $0.266Cached $0.0027

    Xiaomi

    xiaomi/mimo-v2.5
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    41 tokens/s
    Latency
    2.8035 s
    Uptime
    97.1612%
    Data collection
    Prompt No Training
    All contexts
    Input $0.14Output $0.28Cached $0.0028

    StreamLake

    xiaomi/mimo-v2.5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    42 tokens/s
    Latency
    1.368 s
    Uptime
    98.3%
    Data collection
    Prompt No Training
    All contexts
    Input $0.168Output $0.336Cached $0.0034

    Novita

    xiaomi/mimo-v2.5
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    36 tokens/s
    Latency
    3.665 s
    Uptime
    99.3772%
    Data collection
    Zero
    All contexts
    Input $0.168Output $0.336Cached $0.0034
  99. model-router

    xiaomi:frontier

    Stable
    @model-router/xiaomi:frontier

    MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro. This model is an alias to `@xiaomi/mimo-v2.5-pro`.

    All contexts
    Input $0.435Output $0.87Cached $0.0036
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    1M
    ThinkingTool CallingModel Router
    Providers & technical details 8 for @model-router/xiaomi:frontier

    Xiaomi:offline

    Disabledmimo-v2.5-pro
    Context
    1,050,000 tokens
    Max output
    131,072 tokens

    Pricing not published.

    GMICloud:offline

    Disabledxiaomi/mimo-v2.5-pro
    Context
    1,050,000 tokens
    Max output
    945,000 tokens
    Throughput
    18 tokens/s
    Latency
    4.206 s
    Uptime
    57.7358%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3045Output $0.609Cached $0.0028

    DeepInfra

    xiaomi/mimo-v2.5-pro
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    29 tokens/s
    Latency
    0.361 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.39Output $1.17Cached $0.078

    DigitalOcean

    xiaomi/mimo-v2.5-pro
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    21 tokens/s
    Latency
    0.893 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.4Output $1.5Cached $0.08

    AtlasCloud

    xiaomi/mimo-v2.5-pro
    Context
    1,024,000 tokens
    Max output
    131,072 tokens
    Throughput
    36 tokens/s
    Latency
    4.947 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.435Output $0.87Cached $0.0036

    Xiaomi

    xiaomi/mimo-v2.5-pro
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    27 tokens/s
    Latency
    2.6565 s
    Uptime
    97.866%
    Data collection
    Prompt No Training
    All contexts
    Input $0.435Output $0.87Cached $0.0036

    Novita

    xiaomi/mimo-v2.5-pro
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    28 tokens/s
    Latency
    5.238 s
    Uptime
    97.992%
    Data collection
    Zero
    All contexts
    Input $0.4802Output $0.9605Cached $0.004

    StreamLake:offline

    Disabledxiaomi/mimo-v2.5-pro
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    27 tokens/s
    Latency
    2.106 s
    Uptime
    79.8319%
    Data collection
    Prompt No Training
    All contexts
    Input $0.522Output $1.044Cached $0.0043
  100. model-router

    xiaomi:mid

    Stable
    @model-router/xiaomi:mid

    MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. This model is an alias to `@xiaomi/mimo-v2.5`.

    All contexts
    Input $0.168Output $0.336Cached $0.003
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingModel Router
    Providers & technical details 6 for @model-router/xiaomi:mid

    Xiaomi:offline

    Disabledmimo-v2.5
    Context
    1,050,000 tokens
    Max output
    131,072 tokens

    Pricing not published.

    GMICloud:offline

    Disabledxiaomi/mimo-v2.5
    Context
    1,050,000 tokens
    Max output
    945,000 tokens
    Throughput
    2 tokens/s
    Latency
    4.4005 s
    Uptime
    66.619%
    Data collection
    Prompt No Training
    All contexts
    Input $0.119Output $0.238Cached $0.0026

    DeepInfra

    xiaomi/mimo-v2.5
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    13 tokens/s
    Latency
    1.766 s
    Uptime
    99.8196%
    Data collection
    Zero
    All contexts
    Input $0.133Output $0.266Cached $0.0027

    Xiaomi

    xiaomi/mimo-v2.5
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    41 tokens/s
    Latency
    2.8035 s
    Uptime
    97.1612%
    Data collection
    Prompt No Training
    All contexts
    Input $0.14Output $0.28Cached $0.0028

    StreamLake

    xiaomi/mimo-v2.5
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    42 tokens/s
    Latency
    1.368 s
    Uptime
    98.3%
    Data collection
    Prompt No Training
    All contexts
    Input $0.168Output $0.336Cached $0.0034

    Novita

    xiaomi/mimo-v2.5
    Context
    1,048,576 tokens
    Max output
    131,072 tokens
    Throughput
    36 tokens/s
    Latency
    3.665 s
    Uptime
    99.3772%
    Data collection
    Zero
    All contexts
    Input $0.168Output $0.336Cached $0.0034
  101. x-ai

    grok-4.20-multi-agent

    Stable
    @x-ai/grok-4.20-multi-agent

    Grok 4.20 Multi-Agent Beta is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows.

    All contexts
    Input $1.25Output $2.5Cached $0.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Highest
    Context up to
    2M
    Image InputThinking
    Providers & technical details 4 for @x-ai/grok-4.20-multi-agent

    xAI (ZDR)

    x-ai/grok-4.20-multi-agent
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Throughput
    120 tokens/s
    Latency
    36.38 s
    Data collection
    Zero
    All contexts
    Input $1.25Output $2.5Cached $0.2

    xAI

    x-ai/grok-4.20-multi-agent
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Throughput
    373 tokens/s
    Latency
    4.585 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $2.5Cached $0.2

    xAI (Priority) (ZDR)

    x-ai/grok-4.20-multi-agent
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Data collection
    Zero
    All contexts
    Input $2.5Output $5Cached $0.4

    xAI (Priority)

    x-ai/grok-4.20-multi-agent
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $2.5Output $5Cached $0.4
  102. x-ai

    grok-4.20-reasoning

    Stable
    @x-ai/grok-4.20-reasoning

    Grok 4.20 Beta is xAI's newest flagship model with industry-leading speed and agentic tool calling capabilities.

    All contexts
    Input $1.25Output $2.5Cached $0.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    2M
    Image InputThinkingTool Calling
    Providers & technical details 5 for @x-ai/grok-4.20-reasoning

    NagaAI:offline

    Disabledgrok-4.20-beta
    Context
    2,048,000 tokens
    Max output
    2,048,000 tokens
    All contexts
    Input $1Output $3

    xAI (ZDR)

    x-ai/grok-4.20
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Throughput
    69 tokens/s
    Latency
    0.519 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.25Output $2.5Cached $0.2

    xAI

    x-ai/grok-4.20
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Throughput
    74 tokens/s
    Latency
    0.586 s
    Uptime
    99.9527%
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $2.5Cached $0.2

    xAI (Priority) (ZDR)

    x-ai/grok-4.20
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Throughput
    66 tokens/s
    Latency
    0.439 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.5Output $5Cached $0.4

    xAI (Priority)

    x-ai/grok-4.20
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $2.5Output $5Cached $0.4
  103. reka

    reka-edge

    Stable
    @reka/reka-edge

    Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs.

    All contexts
    Input $0.1Output $0.1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Lowest
    Context up to
    16.4K
    Image InputVideo Input
    Providers & technical details 1 for @reka/reka-edge

    Reka

    rekaai/reka-edge
    Context
    16,384 tokens
    Max output
    14,745 tokens
    Throughput
    118 tokens/s
    Latency
    0.357 s
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.1
  104. openai

    gpt-5.4-mini

    Stable
    @openai/gpt-5.4-mini

    GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads.

    All contexts
    Input $0.75Output $4.5Cached $0.075
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    400K
    Image InputThinkingTool CallingFile Input
    Providers & technical details 6 for @openai/gpt-5.4-mini

    NagaAI:offline

    Disabledgpt-5.4-mini
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.38Output $2.25

    OpenAI (Flex)

    openai/gpt-5.4-mini
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $0.375Output $2.25Cached $0.0375

    Azure

    openai/gpt-5.4-mini
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    44 tokens/s
    Latency
    2.182 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.75Output $4.5Cached $0.075

    OpenAI

    openai/gpt-5.4-mini
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    69 tokens/s
    Latency
    1.086 s
    Uptime
    99.951%
    Data collection
    Prompt No Training
    All contexts
    Input $0.75Output $4.5Cached $0.075

    Azure (US)

    openai/gpt-5.4-mini
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $0.825Output $4.95Cached $0.0825

    OpenAI (Fast)

    openai/gpt-5.4-mini
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    52 tokens/s
    Latency
    0.83 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.5Output $9Cached $0.15
  105. openai

    gpt-5.4-nano

    Stable
    @openai/gpt-5.4-nano

    GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks.

    All contexts
    Input $0.2Output $1.25Cached $0.02
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Low
    Context up to
    400K
    Image InputThinkingTool CallingFile Input
    Providers & technical details 5 for @openai/gpt-5.4-nano

    NagaAI:offline

    Disabledgpt-5.4-nano
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.1Output $0.63

    OpenAI (Flex)

    openai/gpt-5.4-nano
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    84 tokens/s
    Latency
    0.587 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.1Output $0.625Cached $0.01

    Azure

    openai/gpt-5.4-nano
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    42 tokens/s
    Latency
    1.55 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.2Output $1.25Cached $0.02

    OpenAI

    openai/gpt-5.4-nano
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    67 tokens/s
    Latency
    0.957 s
    Uptime
    99.5984%
    Data collection
    Prompt No Training
    All contexts
    Input $0.2Output $1.25Cached $0.02

    Azure (US)

    openai/gpt-5.4-nano
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    49 tokens/s
    Latency
    1.272 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.22Output $1.375Cached $0.022
  106. z-ai

    glm-5-turbo

    Stable
    @z-ai/glm-5-turbo

    GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios.

    All contexts
    Input $0.9Output $3Cached $0.24
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    202.8K
    ThinkingTool Calling
    Providers & technical details 2 for @z-ai/glm-5-turbo

    NagaAI:offline

    Disabledglm-5-turbo
    Context
    204,800 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.6Output $2

    Z.AI

    z-ai/glm-5-turbo
    Context
    202,752 tokens
    Max output
    131,072 tokens
    Throughput
    19 tokens/s
    Latency
    8.075 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.2Output $4Cached $0.24
  107. x-ai

    grok-4.20

    Stable
    @x-ai/grok-4.20

    Grok 4.20 Beta is xAI's newest flagship model with industry-leading speed and agentic tool calling capabilities.

    All contexts
    Input $1.25Output $2.5Cached $0.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    2M
    Image InputTool Calling
    Providers & technical details 4 for @x-ai/grok-4.20

    xAI (ZDR)

    x-ai/grok-4.20
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Throughput
    69 tokens/s
    Latency
    0.519 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.25Output $2.5Cached $0.2

    xAI

    x-ai/grok-4.20
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Throughput
    74 tokens/s
    Latency
    0.586 s
    Uptime
    99.9527%
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $2.5Cached $0.2

    xAI (Priority) (ZDR)

    x-ai/grok-4.20
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Throughput
    66 tokens/s
    Latency
    0.439 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.5Output $5Cached $0.4

    xAI (Priority)

    x-ai/grok-4.20
    Context
    2,000,000 tokens
    Max output
    1,800,000 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $2.5Output $5Cached $0.4
  108. qwen

    qwen3.5-9b

    Stable
    @qwen/qwen3.5-9b

    Qwen3.5 9B is an efficient multimodal model from the Qwen3.5 family for low-cost reasoning, coding, and visual understanding workloads.

    All contexts
    Input $0.1Output $0.15
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Low
    Context up to
    262.1K
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 6 for @qwen/qwen3.5-9b

    Darkbloom

    qwen/qwen3.5-9b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    37 tokens/s
    Latency
    0.9 s
    Uptime
    99.6933%
    Data collection
    Prompt No Training
    All contexts
    Input $0.08Output $0.13

    SiliconFlow

    qwen/qwen3.5-9b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    18 tokens/s
    Latency
    1.286 s
    Uptime
    99.8225%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.15

    DeepInfra:offline

    Disabledqwen/qwen3.5-9b
    Context
    262,144 tokens
    Max output
    81,920 tokens
    Throughput
    25 tokens/s
    Latency
    0.634 s
    Uptime
    86.7898%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.15

    Venice

    qwen/qwen3.5-9b
    Context
    256,000 tokens
    Max output
    32,768 tokens
    Throughput
    52 tokens/s
    Latency
    0.704 s
    Uptime
    99.5445%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.15

    Parasail:offline

    Disabledqwen/qwen3.5-9b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    52 tokens/s
    Latency
    0.443 s
    Uptime
    80.4444%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.25

    Together

    qwen/qwen3.5-9b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    32 tokens/s
    Latency
    0.481 s
    Uptime
    97%
    Data collection
    Zero
    All contexts
    Input $0.17Output $0.25
  109. openai

    gpt-5.4

    Stable
    @openai/gpt-5.4

    GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system.

    All contexts
    Input $2.75Output $16.5Cached $0.275
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1.1M
    Image InputThinkingTool Calling
    Providers & technical details 8 for @openai/gpt-5.4

    NagaAI:offline

    Disabledgpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $1.25Output $7.5

    OpenAI (Flex)

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    20 tokens/s
    Latency
    1.172 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $7.5Cached $0.125

    Azure

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    53 tokens/s
    Latency
    1.5335 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.5Output $15Cached $0.25

    OpenAI

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    42 tokens/s
    Latency
    0.978 s
    Uptime
    97.6474%
    Data collection
    Prompt No Training
    All contexts
    Input $2.5Output $15Cached $0.25

    Azure (US)

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $2.75Output $16.5Cached $0.275

    Azure (EU)

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $2.75Output $16.5Cached $0.275

    Amazon Bedrock (US)

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $2.75Output $16.5Cached $0.275

    OpenAI (Fast)

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    112 tokens/s
    Latency
    3.857 s
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $30Cached $0.5
  110. openai

    gpt-5.4-pro

    Stable
    @openai/gpt-5.4-pro

    GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks.

    All contexts
    Input $2.75Output $16.5Cached $0.275
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slowest
    Intelligence
    Highest
    Context up to
    1.1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 8 for @openai/gpt-5.4-pro

    NagaAI:offline

    Disabledgpt-5.4
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $1.25Output $7.5

    OpenAI (Flex)

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    20 tokens/s
    Latency
    1.172 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $7.5Cached $0.125

    Azure

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    53 tokens/s
    Latency
    1.5335 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.5Output $15Cached $0.25

    OpenAI

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    42 tokens/s
    Latency
    0.978 s
    Uptime
    97.6474%
    Data collection
    Prompt No Training
    All contexts
    Input $2.5Output $15Cached $0.25

    Azure (US)

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $2.75Output $16.5Cached $0.275

    Azure (EU)

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $2.75Output $16.5Cached $0.275

    Amazon Bedrock (US)

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $2.75Output $16.5Cached $0.275

    OpenAI (Fast)

    openai/gpt-5.4
    Context
    1,050,000 tokens
    Max output
    128,000 tokens
    Throughput
    112 tokens/s
    Latency
    3.857 s
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $30Cached $0.5
  111. inception

    mercury-2

    Stable
    @inception/mercury-2

    The fastest reasoning LLM and Inception most powerful model.

    All contexts
    Input $0.25Output $0.75Cached $0.025
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    128K
    ThinkingTool CallingDiffusion
    Providers & technical details 2 for @inception/mercury-2

    Inception

    inception/mercury-2
    Context
    128,000 tokens
    Max output
    50,000 tokens
    Throughput
    267 tokens/s
    Latency
    0.904 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.25Output $0.75Cached $0.025

    Inception

    mercury-2
    Context
    128,000 tokens
    Max output
    50,000 tokens
    All contexts
    Input $0.25Output $0.75Cached $0.025
  112. bytedance

    seed-2.0-mini

    Stable
    @bytedance/seed-2.0-mini

    Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment.

    All contexts
    Input $0.1Output $0.4
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Medium
    Context up to
    262.1K
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 1 for @bytedance/seed-2.0-mini

    Seed

    bytedance-seed/seed-2.0-mini
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    58 tokens/s
    Latency
    0.2575 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.4
  113. qwen

    qwen3.5-27b

    Stable
    @qwen/qwen3.5-27b

    The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B.

    All contexts
    Input $0.3Output $2.4
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 6 for @qwen/qwen3.5-27b

    Alibaba

    qwen/qwen3.5-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    23 tokens/s
    Latency
    0.704 s
    Uptime
    99.2604%
    Data collection
    Prompt No Training
    All contexts
    Input $0.195Output $1.56

    SiliconFlow

    qwen/qwen3.5-27b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    11 tokens/s
    Latency
    1.135 s
    Uptime
    99.8092%
    Data collection
    Zero
    All contexts
    Input $0.25Output $2

    DeepInfra

    qwen/qwen3.5-27b
    Context
    262,144 tokens
    Max output
    81,920 tokens
    Throughput
    36 tokens/s
    Latency
    0.613 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.26Output $2.6

    AtlasCloud

    qwen/qwen3.5-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    14 tokens/s
    Latency
    0.9115 s
    Uptime
    98.7013%
    Data collection
    Prompt No Training
    All contexts
    Input $0.27Output $2.16Cached $0.27

    Phala

    qwen/qwen3.5-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    8 tokens/s
    Latency
    1.313 s
    Uptime
    99.25%
    Data collection
    Zero
    All contexts
    Input $0.3Output $2.4Cached $0.15

    Novita

    qwen/qwen3.5-27b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    13 tokens/s
    Latency
    1.0225 s
    Uptime
    99.811%
    Data collection
    Zero
    All contexts
    Input $0.3Output $2.4
  114. qwen

    qwen3.5-flash

    Stable
    @qwen/qwen3.5-flash

    The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.

    All contexts
    Input $0.065Output $0.26
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    1M
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 1 for @qwen/qwen3.5-flash

    Alibaba

    qwen/qwen3.5-flash-02-23
    Context
    1,000,000 tokens
    Max output
    65,536 tokens
    Throughput
    41 tokens/s
    Latency
    0.507 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.065Output $0.26
  115. google

    gemini-3.1-pro

    PreviewStable
    @google/gemini-3.1-pro

    Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows.

    All contexts
    Input $2Output $12Cached $0.2Audio input $2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 7 for @google/gemini-3.1-pro

    NagaAI:offline

    Disabledgemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    All contexts
    Input $1Output $6

    Google (Flex):offline

    Disabledgoogle/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Zero
    All contexts
    Input $1Output $6Cached $0.1Audio input $1

    Google AI Studio (Flex)

    google/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    62 tokens/s
    Latency
    2.767 s
    Data collection
    Prompt No Training
    All contexts
    Input $1Output $6Cached $0.1Audio input $1

    Google

    google/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    83 tokens/s
    Latency
    2.7025 s
    Uptime
    97.3863%
    Data collection
    Zero
    All contexts
    Input $2Output $12Cached $0.2Audio input $2

    Google AI Studio

    google/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    96 tokens/s
    Latency
    3.135 s
    Uptime
    99.7345%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $12Cached $0.2Audio input $2

    Google (Priority)

    google/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Zero
    All contexts
    Input $3.6Output $21.6Cached $0.36Audio input $3.6

    Google AI Studio (Priority)

    google/gemini-3.1-pro-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $3.6Output $21.6Cached $0.36Audio input $3.6
  116. anthropic

    claude-4.6-sonnet

    Stable
    @anthropic/claude-4.6-sonnet

    Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work.

    All contexts
    Input $3Output $15Cached $0.3
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 10 for @anthropic/claude-4.6-sonnet

    NagaAI:offline

    Disabledclaude-sonnet-4.6
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    All contexts
    Input $1.5Output $7.5

    Claude Platform on AWS

    anthropic/claude-sonnet-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    42 tokens/s
    Latency
    0.929 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $3Output $15Cached $0.3

    Azure

    anthropic/claude-sonnet-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    32 tokens/s
    Latency
    1.863 s
    Data collection
    Unknown
    All contexts
    Input $3Output $15Cached $0.3

    Anthropic

    anthropic/claude-sonnet-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    36 tokens/s
    Latency
    1.166 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $3Output $15Cached $0.3

    Amazon Bedrock

    anthropic/claude-sonnet-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    48 tokens/s
    Latency
    1.0775 s
    Uptime
    99.9592%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Google

    anthropic/claude-sonnet-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    43 tokens/s
    Latency
    1.685 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Amazon Bedrock (US)

    anthropic/claude-sonnet-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $3.3Output $16.5Cached $0.33

    Amazon Bedrock (EU)

    anthropic/claude-sonnet-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $3.3Output $16.5Cached $0.33

    Google (EU)

    anthropic/claude-sonnet-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    35 tokens/s
    Latency
    4.7255 s
    Data collection
    Zero
    All contexts
    Input $3.3Output $16.5Cached $0.33

    Google (US)

    anthropic/claude-sonnet-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    12 tokens/s
    Latency
    1.3025 s
    Data collection
    Zero
    All contexts
    Input $3.3Output $16.5Cached $0.33
  117. qwen

    qwen3.5-397b-a17b

    Stable
    @qwen/qwen3.5-397b-a17b

    The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.

    All contexts
    Input $0.575Output $3.6
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 11 for @qwen/qwen3.5-397b-a17b

    NagaAI:offline

    Disabledqwen3.5-397b-a17b
    Context
    268,288 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.2Output $1.17

    Alibaba

    qwen/qwen3.5-397b-a17b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    30 tokens/s
    Latency
    3.875 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.39Output $2.34

    DeepInfra

    qwen/qwen3.5-397b-a17b
    Context
    262,144 tokens
    Max output
    81,920 tokens
    Throughput
    44 tokens/s
    Latency
    0.54 s
    Uptime
    99.904%
    Data collection
    Zero
    All contexts
    Input $0.45Output $3Cached $0.22

    Parasail

    qwen/qwen3.5-397b-a17b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    73 tokens/s
    Latency
    0.353 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.5Output $3.6Cached $0.3

    DigitalOcean

    qwen/qwen3.5-397b-a17b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    9 tokens/s
    Latency
    0.5005 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.55Output $3.5Cached $0.111

    Phala

    qwen/qwen3.5-397b-a17b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    15.5 tokens/s
    Latency
    1.059 s
    Uptime
    99.7549%
    Data collection
    Zero
    All contexts
    Input $0.55Output $3.5Cached $0.225

    AtlasCloud

    qwen/qwen3.5-397b-a17b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    21 tokens/s
    Latency
    1.169 s
    Uptime
    97.3244%
    Data collection
    Prompt No Training
    All contexts
    Input $0.55Output $3.5Cached $0.55

    StreamLake

    qwen/qwen3.5-397b-a17b
    Context
    256,000 tokens
    Max output
    64,000 tokens
    Throughput
    25 tokens/s
    Latency
    0.742 s
    Uptime
    99.1597%
    Data collection
    Prompt No Training
    All contexts
    Input $0.6Output $3.6Cached $0.12

    GMICloud

    qwen/qwen3.5-397b-a17b
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    95 tokens/s
    Latency
    1.3915 s
    Uptime
    99.6633%
    Data collection
    Prompt No Training
    All contexts
    Input $0.6Output $3.6

    Novita

    qwen/qwen3.5-397b-a17b
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    16 tokens/s
    Latency
    6.464 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.6Output $3.6

    Venice

    qwen/qwen3.5-397b-a17b
    Context
    128,000 tokens
    Max output
    32,768 tokens
    Throughput
    46 tokens/s
    Latency
    0.5115 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.75Output $4.5
  118. qwen

    qwen3.5-plus

    Stable
    @qwen/qwen3.5-plus

    The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency.

    All contexts
    Input $0.26Output $1.56
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 1 for @qwen/qwen3.5-plus

    Alibaba

    qwen/qwen3.5-plus-02-15
    Context
    1,000,000 tokens
    Max output
    65,536 tokens
    Throughput
    42 tokens/s
    Latency
    0.9735 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.26Output $1.56
  119. minimax

    m2.5

    Stable
    @minimax/m2.5

    MiniMax-M2.5 is a state-of-the-art large language model built for real-world productivity.

    All contexts
    Input $0.3Output $1.2Cached $0.06
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    204.8K
    ThinkingTool Calling
    Providers & technical details 9 for @minimax/m2.5

    NagaAI:offline

    Disabledminimax-m2.5
    Context
    204,800 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.14Output $0.54

    Venice

    minimax/minimax-m2.5
    Context
    198,000 tokens
    Max output
    32,768 tokens
    Throughput
    22 tokens/s
    Latency
    1.268 s
    Data collection
    Zero
    All contexts
    Input $0.27Output $0.95Cached $0.03

    StreamLake

    minimax/minimax-m2.5
    Context
    200,000 tokens
    Max output
    128,000 tokens
    Throughput
    57 tokens/s
    Latency
    0.863 s
    Uptime
    99.7157%
    Data collection
    Prompt No Training
    All contexts
    Input $0.27Output $1.08Cached $0.027

    AtlasCloud

    minimax/minimax-m2.5
    Context
    196,608 tokens
    Max output
    176,947 tokens
    Throughput
    42.5 tokens/s
    Latency
    2.912 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.295Output $1.2Cached $0.06

    DigitalOcean

    minimax/minimax-m2.5
    Context
    65,536 tokens
    Max output
    58,982 tokens
    Throughput
    64 tokens/s
    Latency
    0.515 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Friendli

    minimax/minimax-m2.5
    Context
    196,608 tokens
    Max output
    176,947 tokens
    Throughput
    112.5 tokens/s
    Latency
    0.4575 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Novita

    minimax/minimax-m2.5
    Context
    204,800 tokens
    Max output
    131,100 tokens
    Throughput
    44 tokens/s
    Latency
    1.1005 s
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.03

    Minimax

    minimax/minimax-m2.5
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    74 tokens/s
    Latency
    0.719 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.03

    Minimax

    minimax/minimax-m2.5
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    73 tokens/s
    Latency
    0.9895 s
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.4Cached $0.06
  120. minimax

    m2.7

    Stable
    @minimax/m2.7

    MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement.

    All contexts
    Input $0.3Output $1.2Cached $0.06
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    204.8K
    ThinkingTool Calling
    Providers & technical details 9 for @minimax/m2.7

    NagaAI:offline

    Disabledminimax-m2.5
    Context
    204,800 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.14Output $0.54

    Venice

    minimax/minimax-m2.5
    Context
    198,000 tokens
    Max output
    32,768 tokens
    Throughput
    22 tokens/s
    Latency
    1.268 s
    Data collection
    Zero
    All contexts
    Input $0.27Output $0.95Cached $0.03

    StreamLake

    minimax/minimax-m2.5
    Context
    200,000 tokens
    Max output
    128,000 tokens
    Throughput
    57 tokens/s
    Latency
    0.863 s
    Uptime
    99.7157%
    Data collection
    Prompt No Training
    All contexts
    Input $0.27Output $1.08Cached $0.027

    AtlasCloud

    minimax/minimax-m2.5
    Context
    196,608 tokens
    Max output
    176,947 tokens
    Throughput
    42.5 tokens/s
    Latency
    2.912 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.295Output $1.2Cached $0.06

    DigitalOcean

    minimax/minimax-m2.5
    Context
    65,536 tokens
    Max output
    58,982 tokens
    Throughput
    64 tokens/s
    Latency
    0.515 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Friendli

    minimax/minimax-m2.5
    Context
    196,608 tokens
    Max output
    176,947 tokens
    Throughput
    112.5 tokens/s
    Latency
    0.4575 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $1.2Cached $0.06

    Novita

    minimax/minimax-m2.5
    Context
    204,800 tokens
    Max output
    131,100 tokens
    Throughput
    44 tokens/s
    Latency
    1.1005 s
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.03

    Minimax

    minimax/minimax-m2.5
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    74 tokens/s
    Latency
    0.719 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.03

    Minimax

    minimax/minimax-m2.5
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    73 tokens/s
    Latency
    0.9895 s
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.4Cached $0.06
  121. z-ai

    glm-5

    Stable
    @z-ai/glm-5

    GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows.

    All contexts
    Input $1Output $3.2Cached $0.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    204.8K
    ThinkingTool Calling
    Providers & technical details 9 for @z-ai/glm-5

    NagaAI:offline

    Disabledglm-5
    Context
    204,800 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.3Output $0.96

    StreamLake

    z-ai/glm-5
    Context
    198,000 tokens
    Max output
    128,000 tokens
    Throughput
    41 tokens/s
    Latency
    2.571 s
    Uptime
    99.7355%
    Data collection
    Prompt No Training
    All contexts
    Input $0.6Output $1.92Cached $0.12

    GMICloud

    z-ai/glm-5
    Context
    202,752 tokens
    Max output
    182,476 tokens
    Throughput
    44 tokens/s
    Latency
    1.44 s
    Uptime
    99.8489%
    Data collection
    Prompt No Training
    All contexts
    Input $0.6Output $1.92Cached $0.12

    Baidu

    z-ai/glm-5
    Context
    202,752 tokens
    Max output
    131,072 tokens
    Throughput
    56 tokens/s
    Latency
    0.839 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.7Output $2.24Cached $0.14

    SiliconFlow

    z-ai/glm-5
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    28 tokens/s
    Latency
    2.419 s
    Uptime
    99.7917%
    Data collection
    Zero
    All contexts
    Input $0.95Output $2.55Cached $0.2

    Amazon Bedrock

    z-ai/glm-5
    Context
    202,752 tokens
    Max output
    131,072 tokens
    Throughput
    41 tokens/s
    Latency
    1.802 s
    Data collection
    Zero
    All contexts
    Input $1Output $3.2

    Venice

    z-ai/glm-5
    Context
    198,000 tokens
    Max output
    32,000 tokens
    Throughput
    51 tokens/s
    Latency
    1.23 s
    Uptime
    97.5124%
    Data collection
    Zero
    All contexts
    Input $1Output $3.2Cached $0.2

    Novita

    z-ai/glm-5
    Context
    202,800 tokens
    Max output
    131,072 tokens
    Throughput
    49 tokens/s
    Latency
    1.138 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1Output $3.2Cached $0.2

    Z.AI

    z-ai/glm-5
    Context
    202,752 tokens
    Max output
    131,072 tokens
    Throughput
    60 tokens/s
    Latency
    6.872 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1Output $3.2Cached $0.2
  122. anthropic

    claude-4.6-opus

    Stable
    @anthropic/claude-4.6-opus

    Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks.

    All contexts
    Input $5Output $25Cached $0.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 7 for @anthropic/claude-4.6-opus

    NagaAI:offline

    Disabledclaude-opus-4.6
    Context
    1,000,000 tokens
    Max output
    131,072 tokens
    All contexts
    Input $2.5Output $12.5

    Claude Platform on AWS

    anthropic/claude-opus-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    36 tokens/s
    Latency
    1.5245 s
    Uptime
    99.9663%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Azure

    anthropic/claude-opus-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Unknown
    All contexts
    Input $5Output $25Cached $0.5

    Google

    anthropic/claude-opus-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    5 tokens/s
    Latency
    1.765 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Amazon Bedrock

    anthropic/claude-opus-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    38.5 tokens/s
    Latency
    1.583 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Anthropic

    anthropic/claude-opus-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Throughput
    31 tokens/s
    Latency
    1.737 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Google (EU)

    anthropic/claude-opus-4.6
    Context
    1,000,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55
  123. moonshotai

    kimi-k2.5

    Stable
    @moonshotai/kimi-k2.5

    Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm.

    All contexts
    Input $0.6Output $3
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    262.1K
    Image InputTool Calling
    Providers & technical details 7 for @moonshotai/kimi-k2.5

    NagaAI:offline

    Disabledkimi-k2.5
    Context
    262,144 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.22Output $1.13

    SiliconFlow

    moonshotai/kimi-k2.5
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    23 tokens/s
    Latency
    1.4325 s
    Uptime
    99.8435%
    Data collection
    Zero
    All contexts
    Input $0.45Output $2.25Cached $0.07

    AtlasCloud

    moonshotai/kimi-k2.5
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    37 tokens/s
    Latency
    0.744 s
    Uptime
    99.3098%
    Data collection
    Prompt No Training
    All contexts
    Input $0.49Output $2.5Cached $0.2

    Venice

    moonshotai/kimi-k2.5
    Context
    256,000 tokens
    Max output
    65,536 tokens
    Throughput
    137 tokens/s
    Latency
    1.118 s
    Uptime
    99.8428%
    Data collection
    Zero
    All contexts
    Input $0.532Output $3.325Cached $0.209

    Novita

    moonshotai/kimi-k2.5
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    33 tokens/s
    Latency
    2.006 s
    Uptime
    98.4615%
    Data collection
    Zero
    All contexts
    Input $0.57Output $2.85Cached $0.095

    Amazon Bedrock (US)

    moonshotai/kimi-k2.5
    Context
    262,144 tokens
    Max output
    131,072 tokens
    Throughput
    31.5 tokens/s
    Latency
    2.0055 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.6Output $3

    Phala

    moonshotai/kimi-k2.5
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    47 tokens/s
    Latency
    1.558 s
    Uptime
    99.1247%
    Data collection
    Zero
    All contexts
    Input $0.6Output $3Cached $0.22
  124. z-ai

    glm-4.7-flash

    Stable
    @z-ai/glm-4.7-flash

    As a SOTA 30B-class model, GLM-4.7-Flash provides a new option that balances efficiency and performance.

    All contexts
    Input $0.0551Output $0.4Cached $0.01
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    131.1K
    ThinkingTool Calling
    Providers & technical details 4 for @z-ai/glm-4.7-flash

    NagaAI:offline

    Disabledglm-4.7-flash
    Context
    204,800 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.03Output $0.2

    Venice

    z-ai/glm-4.7-flash
    Context
    128,000 tokens
    Max output
    16,384 tokens
    Throughput
    16 tokens/s
    Latency
    1.606 s
    Uptime
    99.6218%
    Data collection
    Zero
    All contexts
    Input $0.06Output $0.4Cached $0.01

    Cloudflare

    z-ai/glm-4.7-flash
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    46 tokens/s
    Latency
    0.32 s
    Uptime
    99.04%
    Data collection
    Prompt No Training
    All contexts
    Input $0.0605Output $0.4

    Novita:offline

    Disabledz-ai/glm-4.7-flash
    Context
    200,000 tokens
    Max output
    128,000 tokens
    Throughput
    47 tokens/s
    Latency
    0.921 s
    Uptime
    67.0341%
    Data collection
    Zero
    All contexts
    Input $0.07Output $0.4Cached $0.01
  125. openai

    gpt-5.2-codex

    Stable
    @openai/gpt-5.2-codex

    GPT-5.2-Codex is an enhanced version of GPT-5.1-Codex, optimized for software engineering and coding tasks.

    All contexts
    Input $1.315Output $10.5Cached $0.175
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    400K
    Image InputThinkingTool Calling
    Providers & technical details 2 for @openai/gpt-5.2-codex

    NagaAI:offline

    Disabledgpt-5.2-codex
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.88Output $7

    Azure

    openai/gpt-5.2-codex
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    31 tokens/s
    Latency
    3.214 s
    Data collection
    Zero
    All contexts
    Input $1.75Output $14Cached $0.175
  126. openai

    gpt-5.3-codex

    Stable
    @openai/gpt-5.3-codex

    GPT-5.3-Codex is OpenAI’s most advanced agentic coding model. It pairs the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2.

    All contexts
    Input $1.75Output $14Cached $0.175
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    400K
    Image InputThinkingTool Calling
    Providers & technical details 4 for @openai/gpt-5.3-codex

    NagaAI:offline

    Disabledgpt-5.3-codex
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.88Output $7

    Azure

    openai/gpt-5.3-codex
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    40.5 tokens/s
    Latency
    7.054 s
    Data collection
    Zero
    All contexts
    Input $1.75Output $14Cached $0.175

    OpenAI

    openai/gpt-5.3-codex
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    113 tokens/s
    Latency
    2.281 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.75Output $14Cached $0.175

    OpenAI (Fast)

    openai/gpt-5.3-codex
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $3.5Output $28Cached $0.35
  127. minimax

    m2.1

    Stable
    @minimax/m2.1

    MiniMax-M2.1 is a cutting-edge, lightweight large language model designed for coding, agentic workflows, and modern application development.

    All contexts
    Input $0.3Output $1.2Cached $0.03
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    204.8K
    ThinkingTool Calling
    Providers & technical details 4 for @minimax/m2.1

    NagaAI:offline

    Disabledminimax-m2.1
    Context
    204,800 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.15Output $0.6

    Novita

    minimax/minimax-m2.1
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    31 tokens/s
    Latency
    1.779 s
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.03

    Minimax

    minimax/minimax-m2.1
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    75.5 tokens/s
    Latency
    2.056 s
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.03

    Minimax

    minimax/minimax-m2.1
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    36 tokens/s
    Latency
    1.2145 s
    Data collection
    Zero
    All contexts
    Input $0.3Output $2.4Cached $0.03
  128. z-ai

    glm-4.7

    Stable
    @z-ai/glm-4.7

    GLM-4.7 is Z.AI’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution.

    All contexts
    Input $0.6Output $2.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    204.8K
    ThinkingTool Calling
    Providers & technical details 8 for @z-ai/glm-4.7

    NagaAI:offline

    Disabledglm-4.7
    Context
    204,800 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.2Output $0.88

    DeepInfra

    z-ai/glm-4.7
    Context
    202,752 tokens
    Max output
    131,072 tokens
    Throughput
    17 tokens/s
    Latency
    1.197 s
    Uptime
    98.7159%
    Data collection
    Zero
    All contexts
    Input $0.4Output $1.75Cached $0.08

    Venice

    z-ai/glm-4.7
    Context
    198,000 tokens
    Max output
    16,384 tokens
    Throughput
    15 tokens/s
    Latency
    1.464 s
    Uptime
    98.212%
    Data collection
    Zero
    All contexts
    Input $0.4004Output $1.9292Cached $0.0801

    AtlasCloud:offline

    Disabledz-ai/glm-4.7
    Context
    202,752 tokens
    Max output
    182,476 tokens
    Throughput
    52 tokens/s
    Latency
    1.5925 s
    Uptime
    67.1533%
    Data collection
    Prompt No Training
    All contexts
    Input $0.52Output $1.85Cached $0.12

    Novita

    z-ai/glm-4.7
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    30 tokens/s
    Latency
    2.726 s
    Uptime
    96.5463%
    Data collection
    Zero
    All contexts
    Input $0.54Output $1.98Cached $0.099

    Google

    z-ai/glm-4.7
    Context
    200,000 tokens
    Max output
    128,000 tokens
    Throughput
    25 tokens/s
    Latency
    0.606 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.2

    Z.AI

    z-ai/glm-4.7
    Context
    202,752 tokens
    Max output
    131,072 tokens
    Throughput
    30 tokens/s
    Latency
    12.458 s
    Uptime
    95.6522%
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.2Cached $0.11

    Mancer 2:offline

    Disabledz-ai/glm-4.7
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    14 tokens/s
    Latency
    1.5045 s
    Uptime
    94.8187%
    Data collection
    Zero
    All contexts
    Input $0.7Output $2.5
  129. google

    gemini-3-flash

    PreviewStable
    @google/gemini-3-flash

    Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.

    All contexts
    Input $0.5Output $3Cached $0.05Audio input $1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 7 for @google/gemini-3-flash

    NagaAI:offline

    Disabledgemini-3-flash-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.25Output $1.5

    Google (Flex):offline

    Disabledgoogle/gemini-3-flash-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    29 tokens/s
    Latency
    18.7255 s
    Uptime
    38.2716%
    Data collection
    Zero
    All contexts
    Input $0.25Output $1.5Cached $0.025Audio input $0.5

    Google AI Studio (Flex)

    google/gemini-3-flash-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    37 tokens/s
    Latency
    2.3085 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.25Output $1.5Cached $0.025Audio input $0.5

    Google AI Studio

    google/gemini-3-flash-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    69 tokens/s
    Latency
    1.305 s
    Uptime
    99.9379%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5Output $3Cached $0.05Audio input $1

    Google

    google/gemini-3-flash-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    72 tokens/s
    Latency
    1.175 s
    Uptime
    99.7606%
    Data collection
    Zero
    All contexts
    Input $0.5Output $3Cached $0.05Audio input $1

    Google (Priority)

    google/gemini-3-flash-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    45 tokens/s
    Latency
    1.516 s
    Uptime
    99.9468%
    Data collection
    Zero
    All contexts
    Input $0.9Output $5.4Cached $0.09Audio input $1.8

    Google AI Studio (Priority)

    google/gemini-3-flash-preview
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    104.5 tokens/s
    Latency
    0.946 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.9Output $5.4Cached $0.09Audio input $1.8
  130. openai

    gpt-5.2-pro

    Offline
    @openai/gpt-5.2-pro

    GPT-5.2 Pro is OpenAI's most advanced model, featuring significant upgrades in agentic coding and long-context capabilities compared to GPT-5 Pro.

    All contexts
    Input $11Output $84
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slowest
    Intelligence
    Highest
    Context up to
    Not published
    Image InputThinkingTool Calling
    Providers & technical details 1 for @openai/gpt-5.2-pro

    NagaAI:offline

    Disabledgpt-5.2-pro-2025-12-11
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $11Output $84
  131. openai

    gpt-5.2

    Stable
    @openai/gpt-5.2

    GPT-5.2 is the newest frontier-level model in the GPT-5 line, providing enhanced agentic abilities and better long-context performance than GPT-5.1.

    All contexts
    Input $1.75Output $14Cached $0.175
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Highest
    Context up to
    400K
    Image InputThinkingTool Calling
    Providers & technical details 5 for @openai/gpt-5.2

    NagaAI:offline

    Disabledgpt-5.2-2025-12-11
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.88Output $7

    OpenAI (Flex)

    openai/gpt-5.2
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $0.875Output $7Cached $0.0875

    Azure

    openai/gpt-5.2
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    31 tokens/s
    Latency
    2.094 s
    Uptime
    99.8302%
    Data collection
    Zero
    All contexts
    Input $1.75Output $14Cached $0.175

    OpenAI

    openai/gpt-5.2
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    59 tokens/s
    Latency
    2.516 s
    Uptime
    99.4352%
    Data collection
    Prompt No Training
    All contexts
    Input $1.75Output $14Cached $0.175

    OpenAI (Fast)

    openai/gpt-5.2
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    96.5 tokens/s
    Latency
    2.5075 s
    Data collection
    Prompt No Training
    All contexts
    Input $3.5Output $28Cached $0.35
  132. openai

    gpt-5.2-chat

    DeprecatingStable
    @openai/gpt-5.2-chat

    GPT-5.2 Chat (also known as Instant) is the fast and lightweight version of the 5.2 family, built for low-latency chatting while maintaining strong general intelligence.

    All contexts
    Input $1.315Output $10.5Cached $0.175
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Highest
    Context up to
    128K
    Image InputThinkingTool Calling
    Providers & technical details 2 for @openai/gpt-5.2-chat

    NagaAI:offline

    Disabledgpt-5.2-chat
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.88Output $7

    Azure

    openai/gpt-5.2-chat
    Context
    128,000 tokens
    Max output
    32,000 tokens
    Data collection
    Zero
    All contexts
    Input $1.75Output $14Cached $0.175
  133. openai

    gpt-5.1-codex-max

    Stable
    @openai/gpt-5.1-codex-max

    GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks.

    All contexts
    Input $0.94Output $7.5Cached $0.125
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    400K
    Image InputThinkingTool Calling
    Providers & technical details 2 for @openai/gpt-5.1-codex-max

    NagaAI:offline

    Disabledgpt-5.1-codex-max
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.63Output $5

    Azure

    openai/gpt-5.1-codex-max
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    39.5 tokens/s
    Latency
    2.3735 s
    Data collection
    Zero
    All contexts
    Input $1.25Output $10Cached $0.125
  134. deepseek

    v3.2

    Stable
    @deepseek/v3.2

    DeepSeek-V3.2 is a large language model optimized for high computational efficiency and strong tool-use reasoning.

    All contexts
    Input $1.63Output $1.7667
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    163.8K
    ThinkingTool Calling
    Providers & technical details 16 for @deepseek/v3.2

    NagaAI:offline

    Disableddeepseek-v3.2
    Context
    166,912 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.13Output $0.19

    GMICloud

    deepseek/deepseek-v3.2
    Context
    163,840 tokens
    Max output
    147,456 tokens
    Throughput
    32 tokens/s
    Latency
    2.231 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.2088Output $0.3096Cached $0.0216

    StreamLake

    deepseek/deepseek-v3.2
    Context
    128,000 tokens
    Max output
    64,000 tokens
    Throughput
    24 tokens/s
    Latency
    1.4 s
    Uptime
    99.5869%
    Data collection
    Prompt No Training
    All contexts
    Input $0.2145Output $0.3218Cached $0.0215

    DigitalOcean

    deepseek/deepseek-v3.2
    Context
    163,840 tokens
    Max output
    128,000 tokens
    Throughput
    35 tokens/s
    Latency
    1.208 s
    Uptime
    95.6986%
    Data collection
    Zero
    All contexts
    Input $0.25Output $0.8Cached $0.075

    SiliconFlow

    deepseek/deepseek-v3.2
    Context
    163,840 tokens
    Max output
    147,456 tokens
    Throughput
    14 tokens/s
    Latency
    2.662 s
    Uptime
    99.6377%
    Data collection
    Zero
    All contexts
    Input $0.259Output $0.42Cached $0.135

    DeepInfra

    deepseek/deepseek-v3.2
    Context
    163,840 tokens
    Max output
    16,384 tokens
    Throughput
    24 tokens/s
    Latency
    1.3265 s
    Uptime
    99.9545%
    Data collection
    Zero
    All contexts
    Input $0.26Output $0.38Cached $0.13

    AtlasCloud

    deepseek/deepseek-v3.2
    Context
    163,840 tokens
    Max output
    147,456 tokens
    Throughput
    26 tokens/s
    Latency
    1.184 s
    Uptime
    99.464%
    Data collection
    Prompt No Training
    All contexts
    Input $0.26Output $0.38Cached $0.13

    Venice

    deepseek/deepseek-v3.2
    Context
    160,000 tokens
    Max output
    32,768 tokens
    Throughput
    27 tokens/s
    Latency
    1.402 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.2683Output $0.3902Cached $0.1301

    Novita

    deepseek/deepseek-v3.2
    Context
    163,840 tokens
    Max output
    65,536 tokens
    Throughput
    25 tokens/s
    Latency
    1.242 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.269Output $0.4Cached $0.1345

    Baidu

    deepseek/deepseek-v3.2
    Context
    131,072 tokens
    Max output
    65,536 tokens
    Throughput
    45 tokens/s
    Latency
    0.965 s
    Uptime
    99.9983%
    Data collection
    Prompt No Training
    All contexts
    Input $0.28Output $0.42Cached $0.028

    Alibaba

    deepseek/deepseek-v3.2
    Context
    131,072 tokens
    Max output
    65,536 tokens
    Throughput
    39 tokens/s
    Latency
    0.96 s
    Uptime
    99.2293%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3705Output $1.1115Cached $0.0741

    Friendli

    deepseek/deepseek-v3.2
    Context
    163,840 tokens
    Max output
    147,456 tokens
    Throughput
    23 tokens/s
    Latency
    1.174 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5Output $1.5Cached $0.25

    Google

    deepseek/deepseek-v3.2
    Context
    163,840 tokens
    Max output
    65,536 tokens
    Throughput
    35 tokens/s
    Latency
    1.437 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.56Output $1.68

    Phala

    deepseek/deepseek-v3.2
    Context
    163,840 tokens
    Max output
    64,000 tokens
    Throughput
    8 tokens/s
    Latency
    3.3045 s
    Uptime
    99.8547%
    Data collection
    Zero
    All contexts
    Input $1Output $1Cached $0.5

    Mara

    deepseek/deepseek-v3.2
    Context
    32,768 tokens
    Max output
    7,168 tokens
    Data collection
    Zero
    All contexts
    Input $3Output $4.5

    SambaNova

    deepseek/deepseek-v3.2
    Context
    32,768 tokens
    Max output
    7,168 tokens
    Data collection
    Zero
    All contexts
    Input $3Output $4.5
  135. mistral

    large-2512

    Stable
    @mistral/large-2512

    Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total).

    All contexts
    Input $0.5Output $1.5Cached $0.05
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    Tool Calling
    Providers & technical details 4 for @mistral/large-2512

    NagaAI:offline

    Disabledmistral-large-2512
    Context
    268,288 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.25Output $0.75

    Mistral (ZDR)

    mistralai/mistral-large-2512
    Context
    262,144 tokens
    Max output
    209,715 tokens
    Throughput
    43 tokens/s
    Latency
    0.538 s
    Uptime
    99.9123%
    Data collection
    Zero
    All contexts
    Input $0.5Output $1.5Cached $0.05

    Mistral

    mistralai/mistral-large-2512
    Context
    262,144 tokens
    Max output
    209,715 tokens
    Throughput
    43 tokens/s
    Latency
    0.53 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.5Output $1.5Cached $0.05

    Mistral (EU)

    mistralai/mistral-large-2512
    Context
    262,144 tokens
    Max output
    209,715 tokens
    Throughput
    48 tokens/s
    Latency
    0.5705 s
    Data collection
    Zero
    All contexts
    Input $0.55Output $1.65Cached $0.055
  136. anthropic

    claude-4.5-opus

    Stable
    @anthropic/claude-4.5-opus

    Claude Opus 4.5 is Anthropic’s latest reasoning model, developed for advanced software engineering, complex agent workflows, and extended computer tasks.

    All contexts
    Input $5Output $25Cached $0.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    200K
    Image InputThinkingTool CallingFile Input
    Providers & technical details 8 for @anthropic/claude-4.5-opus

    NagaAI:offline

    Disabledclaude-opus-4.5-20251101
    Context
    200,000 tokens
    Max output
    65,536 tokens
    All contexts
    Input $2.5Output $12.5

    NagaAI:offline

    Disabledclaude-opus-4.5-uncensored
    Context
    200,000 tokens
    Max output
    65,536 tokens
    All contexts
    Input $5Output $25Cached $0.5

    Claude Platform on AWS

    anthropic/claude-opus-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    30.5 tokens/s
    Latency
    4.0985 s
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Azure

    anthropic/claude-opus-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    75 tokens/s
    Latency
    1.202 s
    Data collection
    Unknown
    All contexts
    Input $5Output $25Cached $0.5

    Google

    anthropic/claude-opus-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    50 tokens/s
    Latency
    1.681 s
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Amazon Bedrock

    anthropic/claude-opus-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    4 tokens/s
    Latency
    1.39 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $5Output $25Cached $0.5

    Anthropic

    anthropic/claude-opus-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    35 tokens/s
    Latency
    1.0855 s
    Data collection
    Prompt No Training
    All contexts
    Input $5Output $25Cached $0.5

    Amazon Bedrock (EU)

    anthropic/claude-opus-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Data collection
    Zero
    All contexts
    Input $5.5Output $27.5Cached $0.55
  137. openai

    gpt-5.1

    Stable
    @openai/gpt-5.1

    GPT-5.1 is the newest top-tier model in the GPT-5 series, featuring enhanced general reasoning, better instruction following, and a more natural conversational tone compared to GPT-5.

    All contexts
    Input $1.25Output $10Cached $0.1327
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    400K
    Image InputThinkingTool Calling
    Providers & technical details 6 for @openai/gpt-5.1

    NagaAI:offline

    Disabledgpt-5.1-2025-11-13
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.63Output $5

    OpenAI (Flex)

    openai/gpt-5.1
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    1 tokens/s
    Latency
    27.4585 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.625Output $5Cached $0.0625

    Azure

    openai/gpt-5.1
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    56 tokens/s
    Latency
    1.111 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.25Output $10Cached $0.13

    OpenAI

    openai/gpt-5.1
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    47 tokens/s
    Latency
    1.3125 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $10Cached $0.125

    Azure

    openai/gpt-5.1
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $1.375Output $11Cached $0.143

    OpenAI (Fast)

    openai/gpt-5.1
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    35 tokens/s
    Latency
    0.928 s
    Data collection
    Prompt No Training
    All contexts
    Input $2.5Output $20Cached $0.25
  138. openai

    gpt-5.1-codex

    Stable
    @openai/gpt-5.1-codex

    GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows.

    All contexts
    Input $0.94Output $7.5Cached $0.13
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    400K
    Image InputThinkingTool Calling
    Providers & technical details 2 for @openai/gpt-5.1-codex

    NagaAI:offline

    Disabledgpt-5.1-codex
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.63Output $5

    Azure

    openai/gpt-5.1-codex
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    42.5 tokens/s
    Latency
    2.777 s
    Data collection
    Zero
    All contexts
    Input $1.25Output $10Cached $0.13
  139. moonshotai

    kimi-k2-thinking

    Stable
    @moonshotai/kimi-k2-thinking

    Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning.

    All contexts
    Input $0.6Output $2.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Highest
    Context up to
    262.1K
    ThinkingTool Calling
    Providers & technical details 3 for @moonshotai/kimi-k2-thinking

    NagaAI:offline

    Disabledkimi-k2-thinking
    Context
    262,144 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.3Output $1.25

    Google:offline

    Disabledmoonshotai/kimi-k2-thinking
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    178 tokens/s
    Latency
    0.595 s
    Uptime
    90.404%
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.5

    Novita

    moonshotai/kimi-k2-thinking
    Context
    262,144 tokens
    Max output
    100,352 tokens
    Throughput
    39 tokens/s
    Latency
    0.75 s
    Uptime
    99.8504%
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.5Cached $0.15
  140. anthropic

    claude-4.5-haiku

    Stable
    @anthropic/claude-4.5-haiku

    Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, offering near-frontier intelligence with much lower cost and latency than larger Claude models.

    All contexts
    Input $1.05Output $5.25Cached $0.105
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    200K
    Image InputTool Calling
    Providers & technical details 9 for @anthropic/claude-4.5-haiku

    NagaAI:offline

    Disabledclaude-haiku-4.5-20251001
    Context
    200,000 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.5Output $2.5

    Azure

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    115 tokens/s
    Latency
    0.513 s
    Uptime
    100%
    Data collection
    Unknown
    All contexts
    Input $1Output $5Cached $0.1

    Amazon Bedrock

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    52 tokens/s
    Latency
    0.761 s
    Uptime
    99.9772%
    Data collection
    Zero
    All contexts
    Input $1Output $5Cached $0.1

    Google

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    68 tokens/s
    Latency
    0.795 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1Output $5Cached $0.1

    Anthropic

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    58 tokens/s
    Latency
    0.686 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1Output $5Cached $0.1

    Amazon Bedrock (US)

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Data collection
    Zero
    All contexts
    Input $1.1Output $5.5Cached $0.11

    Google (US)

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Data collection
    Zero
    All contexts
    Input $1.1Output $5.5Cached $0.11

    Amazon Bedrock (EU)

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    79 tokens/s
    Latency
    0.7215 s
    Data collection
    Zero
    All contexts
    Input $1.1Output $5.5Cached $0.11

    Google (EU)

    anthropic/claude-haiku-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    78 tokens/s
    Latency
    0.3855 s
    Data collection
    Zero
    All contexts
    Input $1.1Output $5.5Cached $0.11
  141. z-ai

    glm-4.6

    DeprecatingStable
    @z-ai/glm-4.6

    GLM‐4.6 is a high‐capacity LLM with a 200K‐token context window, strong coding and reasoning abilities, and enhanced tool‐use capabilities.

    All contexts
    Input $0.6Output $2.2Cached $0.11
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    204.8K
    ThinkingTool Calling
    Providers & technical details 6 for @z-ai/glm-4.6

    NagaAI:offline

    Disabledglm-4.6
    Context
    204,800 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.21Output $0.88

    Venice

    z-ai/glm-4.6
    Context
    198,000 tokens
    Max output
    16,384 tokens
    Throughput
    49 tokens/s
    Latency
    0.883 s
    Uptime
    99.8885%
    Data collection
    Zero
    All contexts
    Input $0.43Output $1.75Cached $0.08

    DeepInfra

    z-ai/glm-4.6
    Context
    202,752 tokens
    Max output
    131,072 tokens
    Throughput
    32 tokens/s
    Latency
    0.55 s
    Uptime
    99.8253%
    Data collection
    Zero
    All contexts
    Input $0.5Output $2Cached $0.1

    Novita

    z-ai/glm-4.6
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    34 tokens/s
    Latency
    2.379 s
    Uptime
    97.8959%
    Data collection
    Zero
    All contexts
    Input $0.55Output $2.2Cached $0.11

    AtlasCloud:offline

    Disabledz-ai/glm-4.6
    Context
    202,752 tokens
    Max output
    182,476 tokens
    Throughput
    43 tokens/s
    Latency
    0.904 s
    Uptime
    74.1176%
    Data collection
    Prompt No Training
    All contexts
    Input $0.6Output $2.2Cached $0.11

    Z.AI:offline

    Disabledz-ai/glm-4.6
    Context
    202,752 tokens
    Max output
    131,072 tokens
    Throughput
    35 tokens/s
    Latency
    11.682 s
    Uptime
    92.5676%
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.2Cached $0.11
  142. anthropic

    claude-4.5-sonnet

    Stable
    @anthropic/claude-4.5-sonnet

    Claude Sonnet 4.5 is the newest model in the Sonnet series, offering improvements and updates over Sonnet 4.

    All contexts
    Input $3Output $15Cached $0.3
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 8 for @anthropic/claude-4.5-sonnet

    NagaAI:offline

    Disabledclaude-sonnet-4.5-20250929
    Context
    200,000 tokens
    Max output
    65,536 tokens
    All contexts
    Input $1.5Output $7.5

    Claude Platform on AWS

    anthropic/claude-sonnet-4.5
    Context
    1,000,000 tokens
    Max output
    64,000 tokens
    Throughput
    3 tokens/s
    Latency
    1.253 s
    Uptime
    99.9%
    Data collection
    Prompt No Training
    All contexts
    Input $3Output $15Cached $0.3

    Azure

    anthropic/claude-sonnet-4.5
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    52 tokens/s
    Latency
    1.316 s
    Data collection
    Unknown
    All contexts
    Input $3Output $15Cached $0.3

    Amazon Bedrock

    anthropic/claude-sonnet-4.5
    Context
    1,000,000 tokens
    Max output
    64,000 tokens
    Throughput
    38 tokens/s
    Latency
    1.458 s
    Uptime
    99.9325%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Anthropic

    anthropic/claude-sonnet-4.5
    Context
    1,000,000 tokens
    Max output
    64,000 tokens
    Throughput
    37 tokens/s
    Latency
    1.111 s
    Uptime
    99.7996%
    Data collection
    Prompt No Training
    All contexts
    Input $3Output $15Cached $0.3

    Google

    anthropic/claude-sonnet-4.5
    Context
    1,000,000 tokens
    Max output
    64,000 tokens
    Throughput
    34 tokens/s
    Latency
    1.372 s
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Amazon Bedrock (EU)

    anthropic/claude-sonnet-4.5
    Context
    1,000,000 tokens
    Max output
    64,000 tokens
    Throughput
    41 tokens/s
    Latency
    1.2585 s
    Data collection
    Zero
    All contexts
    Input $3.3Output $16.5Cached $0.33

    Google (US)

    anthropic/claude-sonnet-4.5
    Context
    1,000,000 tokens
    Max output
    64,000 tokens
    Data collection
    Zero
    All contexts
    Input $3.3Output $16.5Cached $0.33
  143. qwen

    qwen3-max

    Stable
    @qwen/qwen3-max

    Qwen3-Max improves instruction following, multilingual ability, and tool use; reduced hallucinations.

    All contexts
    Input $0.585Output $2.925Cached $0.156
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    ThinkingTool Calling
    Providers & technical details 2 for @qwen/qwen3-max

    NagaAI:offline

    Disabledqwen3-max
    Context
    268,288 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.39Output $1.95

    Alibaba

    qwen/qwen3-max
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    53 tokens/s
    Latency
    0.7085 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.78Output $3.9Cached $0.156
  144. deepseek

    v3.1-terminus

    DeprecatingStable
    @deepseek/v3.1-terminus

    DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report.

    All contexts
    Input $0.27Output $1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    163.8K
    ThinkingTool Calling
    Providers & technical details 5 for @deepseek/v3.1-terminus

    NagaAI:offline

    Disableddeepseek-chat-v3.1-terminus
    Context
    166,912 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.14Output $0.5

    SiliconFlow

    deepseek/deepseek-v3.1-terminus
    Context
    163,840 tokens
    Max output
    147,456 tokens
    Throughput
    18 tokens/s
    Latency
    1.571 s
    Uptime
    99.6364%
    Data collection
    Zero
    All contexts
    Input $0.27Output $1

    Novita

    deepseek/deepseek-v3.1-terminus
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    24 tokens/s
    Latency
    2.0655 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.27Output $1Cached $0.135

    AtlasCloud

    deepseek/deepseek-v3.1-terminus
    Context
    131,072 tokens
    Max output
    65,536 tokens
    Throughput
    44.5 tokens/s
    Latency
    1.827 s
    Uptime
    99.5868%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $0.95Cached $0.13

    StreamLake

    deepseek/deepseek-v3.1-terminus
    Context
    128,000 tokens
    Max output
    32,000 tokens
    Throughput
    51 tokens/s
    Latency
    1.614 s
    Uptime
    99.6454%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3426Output $1.0284
  145. moonshotai

    kimi-k2

    Stable
    @moonshotai/kimi-k2

    Model with 1tri total parameters, 32bi activated parameters, optimized for agentic intelligence.

    All contexts
    Input $0.45Output $1.875
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    262.1K
    Tool Calling
    Providers & technical details 3 for @moonshotai/kimi-k2

    Groq (Alt route)

    moonshotai/kimi-k2-instruct-0905
    Context
    262,144 tokens
    Max output
    16,384 tokens

    Pricing not published.

    NagaAI:offline

    Disabledkimi-k2-0905
    Context
    262,144 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.3Output $1.25

    Novita

    moonshotai/kimi-k2-0905
    Context
    262,144 tokens
    Max output
    100,352 tokens
    Throughput
    32 tokens/s
    Latency
    0.766 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.6Output $2.5
  146. openai

    gpt-5

    Stable
    @openai/gpt-5

    OpenAI's newest flagship model for coding, reasoning, and agentic tasks across domains.

    All contexts
    Input $1.25Output $10Cached $0.125
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Highest
    Context up to
    400K
    Image InputThinkingTool Calling
    Providers & technical details 4 for @openai/gpt-5

    NagaAI:offline

    Disabledgpt-5-2025-08-07
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.63Output $5

    Azure

    openai/gpt-5
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    62 tokens/s
    Latency
    3.68 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1.25Output $10Cached $0.125

    OpenAI

    openai/gpt-5
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    75 tokens/s
    Latency
    2.923 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $10Cached $0.125

    Azure

    openai/gpt-5
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Data collection
    Zero
    All contexts
    Input $1.375Output $11Cached $0.1375
  147. openai

    gpt-5-mini

    Stable
    @openai/gpt-5-mini

    GPT-5 mini is a faster, more cost-efficient version of GPT-5.

    All contexts
    Input $0.25Output $2Cached $0.0293
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    400K
    Image InputTool Calling
    Providers & technical details 5 for @openai/gpt-5-mini

    NagaAI:offline

    Disabledgpt-5-mini-2025-08-07
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.13Output $1

    OpenAI (Flex)

    openai/gpt-5-mini
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    26 tokens/s
    Latency
    22.561 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.125Output $1Cached $0.0125

    Azure

    openai/gpt-5-mini
    Context
    400,000 tokens
    Max output
    360,000 tokens
    Throughput
    75 tokens/s
    Latency
    3.426 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.25Output $2Cached $0.03

    OpenAI

    openai/gpt-5-mini
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    77 tokens/s
    Latency
    2.895 s
    Uptime
    99.995%
    Data collection
    Prompt No Training
    All contexts
    Input $0.25Output $2Cached $0.025

    Azure

    openai/gpt-5-mini
    Context
    400,000 tokens
    Max output
    360,000 tokens
    Data collection
    Zero
    All contexts
    Input $0.275Output $2.2Cached $0.033
  148. openai

    gpt-5-nano

    Stable
    @openai/gpt-5-nano

    OpenAI's fastest, cheapest version of GPT-5.

    All contexts
    Input $0.05Output $0.4Cached $0.0087
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Medium
    Context up to
    400K
    Image InputTool Calling
    Providers & technical details 5 for @openai/gpt-5-nano

    NagaAI:offline

    Disabledgpt-5-nano-2025-08-07
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.02Output $0.2

    OpenAI (Flex)

    openai/gpt-5-nano
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    111 tokens/s
    Latency
    0.986 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.025Output $0.2Cached $0.0025

    Azure

    openai/gpt-5-nano
    Context
    400,000 tokens
    Max output
    360,000 tokens
    Throughput
    74 tokens/s
    Latency
    1.68 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.05Output $0.4Cached $0.01

    OpenAI

    openai/gpt-5-nano
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    130.5 tokens/s
    Latency
    1.932 s
    Uptime
    98.5678%
    Data collection
    Prompt No Training
    All contexts
    Input $0.05Output $0.4Cached $0.005

    Azure

    openai/gpt-5-nano
    Context
    400,000 tokens
    Max output
    360,000 tokens
    Data collection
    Zero
    All contexts
    Input $0.055Output $0.44Cached $0.011
  149. anthropic

    claude-4.1-opus

    Stable
    @anthropic/claude-4.1-opus

    Claude Opus 4.1 is Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks.

    All contexts
    Input $15Output $75Cached $1.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    200K
    Image InputThinkingTool CallingFile Input
    Providers & technical details 3 for @anthropic/claude-4.1-opus

    NagaAI:offline

    Disabledclaude-opus-4.1-20250805
    Context
    200,000 tokens
    Max output
    65,536 tokens
    All contexts
    Input $7.5Output $37.5

    Amazon Bedrock

    anthropic/claude-opus-4.1
    Context
    200,000 tokens
    Max output
    32,000 tokens
    Throughput
    4 tokens/s
    Latency
    3.014 s
    Data collection
    Zero
    All contexts
    Input $15Output $75Cached $1.5

    Google

    anthropic/claude-opus-4.1
    Context
    200,000 tokens
    Max output
    32,000 tokens
    Throughput
    3 tokens/s
    Latency
    1.9485 s
    Data collection
    Zero
    All contexts
    Input $15Output $75Cached $1.5
  150. openai

    gpt-oss-120b

    Stable
    @openai/gpt-oss-120b

    OpenAI's flagship open source model, built on a Mixture-of-Experts (MoE) architecture with 120 billion parameters and 128 experts.

    All contexts
    Input $0.15Output $0.6
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    131.1K
    ThinkingTool Calling
    Providers & technical details 25 for @openai/gpt-oss-120b

    Groq (Alt route)

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.15Output $0.6Cached $0.075

    NagaAI:offline

    Disabledgpt-oss-120b
    Context
    131,072 tokens
    Max output
    32,768 tokens
    All contexts
    Input $0.02Output $0.08

    AkashML

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    59 tokens/s
    Latency
    0.458 s
    Uptime
    99.9867%
    Data collection
    Zero
    All contexts
    Input $0.03Output $0.17Cached $0.03

    CoreWeave

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    39 tokens/s
    Latency
    0.36 s
    Uptime
    99.9751%
    Data collection
    Zero
    All contexts
    Input $0.03Output $0.17Cached $0.03

    DekaLLM

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    20 tokens/s
    Latency
    0.745 s
    Uptime
    99.7665%
    Data collection
    Prompt No Training
    All contexts
    Input $0.03Output $0.18

    DeepInfra

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    27 tokens/s
    Latency
    0.781 s
    Uptime
    97.879%
    Data collection
    Zero
    All contexts
    Input $0.037Output $0.17

    Novita

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    103 tokens/s
    Latency
    0.805 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.05Output $0.25

    DigitalOcean

    openai/gpt-oss-120b
    Context
    128,000 tokens
    Max output
    115,200 tokens
    Throughput
    41 tokens/s
    Latency
    0.494 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.055Output $0.385Cached $0.02

    Mancer 2

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    131 tokens/s
    Latency
    0.463 s
    Uptime
    99.9006%
    Data collection
    Zero
    All contexts
    Input $0.055Output $0.5

    Google

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    202 tokens/s
    Latency
    0.254 s
    Uptime
    99.9248%
    Data collection
    Zero
    All contexts
    Input $0.09Output $0.36

    BaseTen

    openai/gpt-oss-120b
    Context
    128,072 tokens
    Max output
    115,264 tokens
    Throughput
    239 tokens/s
    Latency
    0.219 s
    Uptime
    99.9603%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.5Cached $0.1

    BaseTen

    openai/gpt-oss-120b
    Context
    128,072 tokens
    Max output
    115,264 tokens
    Throughput
    216 tokens/s
    Latency
    0.254 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.5Cached $0.1

    Parasail

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    156 tokens/s
    Latency
    0.42 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.75Cached $0.055

    SambaNova

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    485 tokens/s
    Latency
    0.351 s
    Uptime
    99.9657%
    Data collection
    Zero
    All contexts
    Input $0.14Output $0.95

    Amazon Bedrock (EU)

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    73 tokens/s
    Latency
    1.045 s
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.6

    Nebius

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    236 tokens/s
    Latency
    0.184 s
    Uptime
    99.9822%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.6

    Amazon Bedrock

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    203 tokens/s
    Latency
    0.413 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.6

    DeepInfra

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    16,384 tokens
    Throughput
    135 tokens/s
    Latency
    0.416 s
    Uptime
    99.9295%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.6

    SiliconFlow

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    8,192 tokens
    Throughput
    42 tokens/s
    Latency
    1.203 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.6Cached $0.075

    Phala

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    105 tokens/s
    Latency
    0.909 s
    Uptime
    99.6689%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.6

    Together

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    141 tokens/s
    Latency
    0.166 s
    Uptime
    99.2308%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.6

    Groq

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    65,536 tokens
    Throughput
    327 tokens/s
    Latency
    0.154 s
    Uptime
    99.9871%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.6Cached $0.075

    Mara:offline

    Disabledopenai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    162 tokens/s
    Latency
    0.928 s
    Uptime
    87.812%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.75

    DeepInfra

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    158 tokens/s
    Latency
    1.394 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.2Output $0.95

    Cerebras

    openai/gpt-oss-120b
    Context
    131,072 tokens
    Max output
    40,960 tokens
    Throughput
    646.5 tokens/s
    Latency
    0.228 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.35Output $0.75Cached $0.35
  151. openai

    gpt-oss-20b

    Stable
    @openai/gpt-oss-20b

    OpenAI's flagship open source model, built on a Mixture-of-Experts (MoE) architecture with 20 billion parameters and 128 experts.

    All contexts
    Input $0.0467Output $0.15
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    Medium
    Context up to
    131.1K
    ThinkingTool Calling
    Providers & technical details 16 for @openai/gpt-oss-20b

    Groq (Alt route)

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.075Output $0.3

    NagaAI:offline

    Disabledgpt-oss-20b
    Context
    131,072 tokens
    Max output
    32,768 tokens
    All contexts
    Input $0.01Output $0.07

    Darkbloom

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    43 tokens/s
    Latency
    3.924 s
    Uptime
    99.5081%
    Data collection
    Prompt No Training
    All contexts
    Input $0.02Output $0.1

    AkashML

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    28 tokens/s
    Latency
    1.301 s
    Uptime
    99.7994%
    Data collection
    Zero
    All contexts
    Input $0.02Output $0.1

    DekaLLM

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    23 tokens/s
    Latency
    0.946 s
    Uptime
    99.9323%
    Data collection
    Prompt No Training
    All contexts
    Input $0.029Output $0.14

    CoreWeave

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    167 tokens/s
    Latency
    0.089 s
    Uptime
    99.9987%
    Data collection
    Zero
    All contexts
    Input $0.03Output $0.13Cached $0.03

    DeepInfra

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    104 tokens/s
    Latency
    0.173 s
    Uptime
    99.994%
    Data collection
    Zero
    All contexts
    Input $0.03Output $0.14

    Parasail

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    104 tokens/s
    Latency
    0.588 s
    Uptime
    99.9901%
    Data collection
    Zero
    All contexts
    Input $0.03Output $0.15Cached $0.02

    Phala

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    53 tokens/s
    Latency
    0.329 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.04Output $0.15

    Novita

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    183 tokens/s
    Latency
    0.599 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.04Output $0.15

    SiliconFlow

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    8,192 tokens
    Throughput
    79 tokens/s
    Latency
    1.053 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.04Output $0.18

    Together

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    80 tokens/s
    Latency
    0.263 s
    Data collection
    Zero
    All contexts
    Input $0.05Output $0.2

    Amazon Bedrock (EU)

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    50 tokens/s
    Latency
    0.316 s
    Data collection
    Zero
    All contexts
    Input $0.07Output $0.15

    Amazon Bedrock

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    405 tokens/s
    Latency
    0.403 s
    Uptime
    99.9038%
    Data collection
    Zero
    All contexts
    Input $0.07Output $0.15

    Google (US)

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    184.5 tokens/s
    Latency
    0.3645 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.07Output $0.25

    Groq

    openai/gpt-oss-20b
    Context
    131,072 tokens
    Max output
    65,536 tokens
    Throughput
    540 tokens/s
    Latency
    0.464 s
    Uptime
    99.7907%
    Data collection
    Zero
    All contexts
    Input $0.075Output $0.3Cached $0.0375
  152. google

    gemini-2.5-flash-lite

    Stable
    @google/gemini-2.5-flash-lite

    A Gemini 2.5 Flash model optimized for cost efficiency and low latency.

    All contexts
    Input $0.1Output $0.4Cached $0.01Audio input $0.3
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    Medium
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool Calling
    Providers & technical details 6 for @google/gemini-2.5-flash-lite

    NagaAI:offline

    Disabledgemini-2.5-flash-lite
    Context
    1,000,000 tokens
    Max output
    64,000 tokens
    All contexts
    Input $0.05Output $0.2

    Google (EU)

    google/gemini-2.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    76 tokens/s
    Latency
    0.473 s
    Uptime
    99.9917%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.4Cached $0.01Audio input $0.3

    Google

    google/gemini-2.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    61 tokens/s
    Latency
    1.09 s
    Uptime
    99.9629%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.4Cached $0.01Audio input $0.3

    Google AI Studio (Flex)

    google/gemini-2.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    51 tokens/s
    Latency
    0.973 s
    Uptime
    99.9873%
    Data collection
    Prompt No Training
    All contexts
    Input $0.05Output $0.2Cached $0.005Audio input $0.15

    Google AI Studio

    google/gemini-2.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    158 tokens/s
    Latency
    0.423 s
    Uptime
    98.1541%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1Output $0.4Cached $0.01Audio input $0.3

    Google AI Studio (Priority)

    google/gemini-2.5-flash-lite
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    144 tokens/s
    Latency
    0.298 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.18Output $0.72Cached $0.018Audio input $0.54
  153. google

    gemini-2.5-flash

    Stable
    @google/gemini-2.5-flash

    Google's best model in terms of price-performance, offering well-rounded capabilities. 2.5 Flash is best for large scale processing, low-latency, high volume tasks that require thinking, and agentic use cases.

    All contexts
    Input $0.3Output $2.5Cached $0.03Audio input $1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 8 for @google/gemini-2.5-flash

    NagaAI:offline

    Disabledgemini-2.5-flash
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.15Output $1.25

    Google (EU)

    google/gemini-2.5-flash
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    98 tokens/s
    Latency
    0.712 s
    Uptime
    99.5643%
    Data collection
    Zero
    All contexts
    Input $0.3Output $2.5Cached $0.03Audio input $1

    Google

    google/gemini-2.5-flash
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    72 tokens/s
    Latency
    0.571 s
    Uptime
    99.8444%
    Data collection
    Zero
    All contexts
    Input $0.3Output $2.5Cached $0.03Audio input $1

    Google (Priority)

    google/gemini-2.5-flash
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    14 tokens/s
    Latency
    0.7795 s
    Data collection
    Zero
    All contexts
    Input $0.54Output $4.5Cached $0.054Audio input $1.8

    Google AI Studio (Flex)

    google/gemini-2.5-flash
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    3 tokens/s
    Latency
    2.385 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $1.25Cached $0.015Audio input $0.5

    Google AI Studio

    google/gemini-2.5-flash
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    63 tokens/s
    Latency
    0.593 s
    Uptime
    99.9849%
    Data collection
    Prompt No Training
    All contexts
    Input $0.3Output $2.5Cached $0.03Audio input $1

    Google:offline

    Disabledgoogle/gemini-2.5-flash
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    83 tokens/s
    Latency
    1.467 s
    Uptime
    78.6813%
    Data collection
    Zero
    All contexts
    Input $0.3Output $2.5Cached $0.03Audio input $1

    Google AI Studio (Priority)

    google/gemini-2.5-flash
    Context
    1,048,576 tokens
    Max output
    65,535 tokens
    Throughput
    141 tokens/s
    Latency
    0.407 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.54Output $4.5Cached $0.054Audio input $1.8
  154. google

    gemini-2.5-pro

    Stable
    @google/gemini-2.5-pro

    One of the most powerful models today.

    All contexts
    Input $1.25Output $10Cached $0.125Audio input $1.25
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Highest
    Context up to
    1M
    Audio InputImage InputVideo InputThinkingTool CallingFile Input
    Providers & technical details 8 for @google/gemini-2.5-pro

    NagaAI:offline

    Disabledgemini-2.5-pro
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.63Output $5

    Google

    google/gemini-2.5-pro
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    81 tokens/s
    Latency
    2.475 s
    Uptime
    99.4053%
    Data collection
    Zero
    All contexts
    Input $1.25Output $10Cached $0.125Audio input $1.25

    Google (Priority)

    google/gemini-2.5-pro
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    94 tokens/s
    Latency
    2.011 s
    Data collection
    Zero
    All contexts
    Input $2.25Output $18Cached $0.225Audio input $2.25

    Google AI Studio (Flex)

    google/gemini-2.5-pro
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    21 tokens/s
    Latency
    4.507 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.625Output $5Cached $0.0625Audio input $0.625

    Google (EU)

    google/gemini-2.5-pro
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    76 tokens/s
    Latency
    3.538 s
    Data collection
    Zero
    All contexts
    Input $1.25Output $10Cached $0.125Audio input $1.25

    Google AI Studio

    google/gemini-2.5-pro
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    77 tokens/s
    Latency
    4.408 s
    Uptime
    95.1417%
    Data collection
    Prompt No Training
    All contexts
    Input $1.25Output $10Cached $0.125Audio input $1.25

    Google (US)

    google/gemini-2.5-pro
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Throughput
    81 tokens/s
    Latency
    17.3695 s
    Data collection
    Zero
    All contexts
    Input $1.25Output $10Cached $0.125Audio input $1.25

    Google AI Studio (Priority)

    google/gemini-2.5-pro
    Context
    1,048,576 tokens
    Max output
    65,536 tokens
    Data collection
    Prompt No Training
    All contexts
    Input $2.25Output $18Cached $0.225Audio input $2.25
  155. deepseek

    r1

    Stable
    @deepseek/r1

    The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528.

    All contexts
    Input $0.5Output $2.0432
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    163.8K
    ThinkingTool Calling
    Providers & technical details 5 for @deepseek/r1

    NagaAI:offline

    Disableddeepseek-reasoner-0528
    Context
    163,840 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.28Output $1.1

    DeepInfra

    deepseek/deepseek-r1-0528
    Context
    163,840 tokens
    Max output
    32,768 tokens
    Throughput
    22 tokens/s
    Latency
    0.6985 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.5Output $2.15Cached $0.35

    SiliconFlow

    deepseek/deepseek-r1-0528
    Context
    163,840 tokens
    Max output
    147,456 tokens
    Throughput
    22 tokens/s
    Latency
    0.926 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.5Output $2.18

    StreamLake

    deepseek/deepseek-r1-0528
    Context
    128,000 tokens
    Max output
    32,000 tokens
    Throughput
    49 tokens/s
    Latency
    3.237 s
    Uptime
    99.6283%
    Data collection
    Prompt No Training
    All contexts
    Input $0.571Output $2.286

    Novita

    deepseek/deepseek-r1-0528
    Context
    163,840 tokens
    Max output
    32,768 tokens
    Throughput
    21 tokens/s
    Latency
    0.832 s
    Data collection
    Zero
    All contexts
    Input $0.7Output $2.5Cached $0.35
  156. anthropic

    claude-4-sonnet

    Stable
    @anthropic/claude-4-sonnet

    Anthropic's mid-size model with superior intelligence for high-volume uses in coding, in-depth research, agents, & more.

    All contexts
    Input $3Output $15Cached $0.3
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    1M
    Image InputThinkingTool CallingFile Input
    Providers & technical details 6 for @anthropic/claude-4-sonnet

    NagaAI:offline

    Disabledclaude-sonnet-4-20250514
    Context
    200,000 tokens
    Max output
    65,536 tokens
    All contexts
    Input $1.5Output $7.5

    Amazon Bedrock (EU)

    anthropic/claude-sonnet-4
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    3 tokens/s
    Latency
    1.79 s
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Google

    anthropic/claude-sonnet-4
    Context
    1,000,000 tokens
    Max output
    64,000 tokens
    Throughput
    41 tokens/s
    Latency
    1.295 s
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Google (EU)

    anthropic/claude-sonnet-4
    Context
    1,000,000 tokens
    Max output
    64,000 tokens
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Amazon Bedrock

    anthropic/claude-sonnet-4
    Context
    200,000 tokens
    Max output
    64,000 tokens
    Throughput
    54 tokens/s
    Latency
    1.4365 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3

    Google

    anthropic/claude-sonnet-4
    Context
    1,000,000 tokens
    Max output
    64,000 tokens
    Data collection
    Zero
    All contexts
    Input $3Output $15Cached $0.3
  157. openai

    o3

    DeprecatingStable
    @openai/o3

    A well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks.

    All contexts
    Input $1.5Output $6Cached $0.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    200K
    Image InputThinkingTool Calling
    Providers & technical details 2 for @openai/o3

    NagaAI:offline

    Disabledo3-2025-04-16
    Context
    200,000 tokens
    Max output
    100,000 tokens
    All contexts
    Input $1Output $4

    OpenAI

    openai/o3
    Context
    200,000 tokens
    Max output
    100,000 tokens
    Throughput
    74 tokens/s
    Latency
    2.2125 s
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $8Cached $0.5
  158. openai

    o4-mini

    DeprecatingStable
    @openai/o4-mini

    Optimized for fast, effective reasoning with exceptionally efficient performance in coding and visual tasks.

    All contexts
    Input $0.825Output $3.3Cached $0.275
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    200K
    Image InputThinkingTool Calling
    Providers & technical details 2 for @openai/o4-mini

    NagaAI:offline

    Disabledo4-mini-2025-04-16
    Context
    200,000 tokens
    Max output
    100,000 tokens
    All contexts
    Input $0.55Output $2.2

    OpenAI

    openai/o4-mini
    Context
    200,000 tokens
    Max output
    100,000 tokens
    Throughput
    129 tokens/s
    Latency
    3.5515 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $1.1Output $4.4Cached $0.275
  159. openai

    gpt-4.1

    DeprecatingStable
    @openai/gpt-4.1

    Versatile, highly intelligent, and top-of-the-line. One of the most capable models currently available.

    All contexts
    Input $2Output $8Cached $0.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Highest
    Context up to
    1M
    Image InputTool Calling
    Providers & technical details 4 for @openai/gpt-4.1

    NagaAI:offline

    Disabledgpt-4.1-2025-04-14
    Context
    1,047,576 tokens
    Max output
    32,768 tokens
    All contexts
    Input $1Output $4

    Azure

    openai/gpt-4.1
    Context
    1,047,576 tokens
    Max output
    942,818 tokens
    Throughput
    59.5 tokens/s
    Latency
    1.188 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2Output $8Cached $0.5

    OpenAI

    openai/gpt-4.1
    Context
    1,047,576 tokens
    Max output
    32,768 tokens
    Throughput
    61 tokens/s
    Latency
    0.833 s
    Uptime
    99.4636%
    Data collection
    Prompt No Training
    All contexts
    Input $2Output $8Cached $0.5

    Azure

    openai/gpt-4.1
    Context
    1,047,576 tokens
    Max output
    942,818 tokens
    Throughput
    17.5 tokens/s
    Latency
    0.588 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $2.2Output $8.8Cached $0.55
  160. openai

    gpt-4.1-mini

    DeprecatingStable
    @openai/gpt-4.1-mini

    Fast and cheap for focused tasks.

    All contexts
    Input $0.4Output $1.6Cached $0.1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Medium
    Context up to
    1M
    Image InputTool Calling
    Providers & technical details 4 for @openai/gpt-4.1-mini

    NagaAI:offline

    Disabledgpt-4.1-mini-2025-04-14
    Context
    1,047,576 tokens
    Max output
    32,768 tokens
    All contexts
    Input $0.2Output $0.8

    Azure

    openai/gpt-4.1-mini
    Context
    1,047,576 tokens
    Max output
    942,818 tokens
    Throughput
    44 tokens/s
    Latency
    1.144 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.4Output $1.6Cached $0.1

    OpenAI

    openai/gpt-4.1-mini
    Context
    1,047,576 tokens
    Max output
    32,768 tokens
    Throughput
    34 tokens/s
    Latency
    0.934 s
    Uptime
    99.7801%
    Data collection
    Prompt No Training
    All contexts
    Input $0.4Output $1.6Cached $0.1

    Azure

    openai/gpt-4.1-mini
    Context
    1,047,576 tokens
    Max output
    942,818 tokens
    Throughput
    87 tokens/s
    Latency
    1.174 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.44Output $1.76Cached $0.11
  161. openai

    gpt-4.1-nano

    DeprecatingStable
    @openai/gpt-4.1-nano

    The fastest and cheapest GPT 4.1 model.

    All contexts
    Input $0.1Output $0.4Cached $0.0293
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Low
    Context up to
    1M
    Image InputTool Calling
    Providers & technical details 4 for @openai/gpt-4.1-nano

    NagaAI:offline

    Disabledgpt-4.1-nano-2025-04-14
    Context
    1,047,576 tokens
    Max output
    32,768 tokens
    All contexts
    Input $0.05Output $0.2

    Azure

    openai/gpt-4.1-nano
    Context
    1,047,576 tokens
    Max output
    942,818 tokens
    Throughput
    46 tokens/s
    Latency
    1.0805 s
    Uptime
    99.8228%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.4Cached $0.03

    OpenAI

    openai/gpt-4.1-nano
    Context
    1,047,576 tokens
    Max output
    32,768 tokens
    Throughput
    90 tokens/s
    Latency
    0.789 s
    Uptime
    99.9754%
    Data collection
    Prompt No Training
    All contexts
    Input $0.1Output $0.4Cached $0.025

    Azure

    openai/gpt-4.1-nano
    Context
    1,047,576 tokens
    Max output
    942,818 tokens
    Data collection
    Zero
    All contexts
    Input $0.11Output $0.44Cached $0.033
  162. cohere

    command-a

    Stable
    @cohere/command-a

    Command A is Cohere's most performant model to date, excelling at tool use, agents, retrieval augmented generation (RAG), and multilingual use cases. Command A has a context length of 256K, only requires two GPUs to run, and has 150% higher throughput compared to Command R+ 08-2024.

    All contexts
    Input $1.875Output $7.5
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    256K
    Image InputTool Calling
    Providers & technical details 2 for @cohere/command-a

    NagaAI:offline

    Disabledcommand-a-03-2025
    Context
    262,144 tokens
    Max output
    65,536 tokens
    All contexts
    Input $1.25Output $5

    Cohere

    cohere/command-a
    Context
    256,000 tokens
    Max output
    8,192 tokens
    Throughput
    15 tokens/s
    Latency
    0.343 s
    Data collection
    Prompt No Training
    All contexts
    Input $2.5Output $10
  163. perplexity

    sonar-pro

    Stable
    @perplexity/sonar-pro

    Sonar Pro is an enterprise-grade API from Perplexity, built for advanced, multi-step queries with added extensibility.

    All contexts
    Input $2.25Output $11.25
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    200K
    Thinking
    Providers & technical details 2 for @perplexity/sonar-pro

    NagaAI:offline

    Disabledsonar-pro
    Context
    131,072 tokens
    Max output
    131,072 tokens
    All contexts
    Input $1.5Output $7.5

    Perplexity

    perplexity/sonar-pro
    Context
    200,000 tokens
    Max output
    8,000 tokens
    Throughput
    71 tokens/s
    Latency
    3.156 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $3Output $15
  164. openai

    o3-mini

    DeprecatingStable
    @openai/o3-mini

    o3-mini provides high intelligence at the same cost and latency targets of previous versions of o-mini series.

    All contexts
    Input $0.825Output $3.3Cached $0.55
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    200K
    ThinkingTool Calling
    Providers & technical details 2 for @openai/o3-mini

    NagaAI:offline

    Disabledo3-mini-2025-01-31
    Context
    200,000 tokens
    Max output
    100,000 tokens
    All contexts
    Input $0.55Output $2.2

    OpenAI

    openai/o3-mini
    Context
    200,000 tokens
    Max output
    100,000 tokens
    Throughput
    157 tokens/s
    Latency
    1.8065 s
    Data collection
    Prompt No Training
    All contexts
    Input $1.1Output $4.4Cached $0.55
  165. perplexity

    sonar

    Stable
    @perplexity/sonar

    Sonar is Perplexity’s lightweight, affordable, and fast question-answering model, now featuring citations and customizable sources.

    All contexts
    Input $0.75Output $0.75
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    127.1K
    Text
    Providers & technical details 2 for @perplexity/sonar

    NagaAI:offline

    Disabledsonar
    Context
    131,072 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.5Output $0.5

    Perplexity

    perplexity/sonar
    Context
    127,072 tokens
    Max output
    114,364 tokens
    Throughput
    60 tokens/s
    Latency
    2.786 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $1Output $1
  166. metaai

    llama-3.3-70b

    DeprecatingStable
    @metaai/llama-3.3-70b

    Previous generation model with many parameters and surprisingly fast speed.

    All contexts
    Input $0.72Output $0.72
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    Medium
    Context up to
    131.1K
    Tool Calling
    Providers & technical details 14 for @metaai/llama-3.3-70b

    Groq (Alt route)

    llama-3.3-70b-versatile
    Context
    131,072 tokens
    Max output
    32,768 tokens

    Pricing not published.

    NagaAI:offline

    Disabledllama-3.3-70b-instruct
    Context
    131,072 tokens
    Max output
    32,768 tokens
    All contexts
    Input $0.35Output $0.35

    DeepInfra

    meta-llama/llama-3.3-70b-instruct
    Context
    131,072 tokens
    Max output
    16,384 tokens
    Throughput
    16 tokens/s
    Latency
    0.389 s
    Uptime
    99.2018%
    Data collection
    Zero
    All contexts
    Input $0.1Output $0.32

    Novita

    meta-llama/llama-3.3-70b-instruct
    Context
    12,288 tokens
    Max output
    11,059 tokens
    Throughput
    37 tokens/s
    Latency
    1.1845 s
    Uptime
    99.8747%
    Data collection
    Zero
    All contexts
    Input $0.135Output $0.4

    AkashML

    meta-llama/llama-3.3-70b-instruct
    Context
    131,072 tokens
    Max output
    128,000 tokens
    Throughput
    32 tokens/s
    Latency
    1.637 s
    Uptime
    99.829%
    Data collection
    Zero
    All contexts
    Input $0.2Output $0.52Cached $0.1

    Parasail

    meta-llama/llama-3.3-70b-instruct
    Context
    131,072 tokens
    Max output
    16,384 tokens
    Throughput
    44 tokens/s
    Latency
    0.742 s
    Uptime
    99.9154%
    Data collection
    Zero
    All contexts
    Input $0.22Output $0.5Cached $0.11

    Crusoe

    meta-llama/llama-3.3-70b-instruct
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    61 tokens/s
    Latency
    0.619 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.25Output $0.75Cached $0.13

    Cloudflare

    meta-llama/llama-3.3-70b-instruct
    Context
    24,000 tokens
    Max output
    21,600 tokens
    Throughput
    35 tokens/s
    Latency
    0.557 s
    Uptime
    98.8686%
    Data collection
    Prompt No Training
    All contexts
    Input $0.293Output $2.253

    SambaNova

    meta-llama/llama-3.3-70b-instruct
    Context
    131,072 tokens
    Max output
    3,072 tokens
    Throughput
    58 tokens/s
    Latency
    0.7605 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.45Output $0.9

    Groq

    meta-llama/llama-3.3-70b-instruct
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    155 tokens/s
    Latency
    0.204 s
    Uptime
    99.9413%
    Data collection
    Zero
    All contexts
    Input $0.59Output $0.79Cached $0.295

    CoreWeave

    meta-llama/llama-3.3-70b-instruct
    Context
    128,000 tokens
    Max output
    115,200 tokens
    Throughput
    69 tokens/s
    Latency
    0.673 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.71Output $0.71Cached $0.71

    Google (US)

    meta-llama/llama-3.3-70b-instruct
    Context
    128,000 tokens
    Max output
    8,192 tokens
    Throughput
    100 tokens/s
    Latency
    1.985 s
    Data collection
    Zero
    All contexts
    Input $0.72Output $0.72

    Google

    meta-llama/llama-3.3-70b-instruct
    Context
    128,000 tokens
    Max output
    115,200 tokens
    Data collection
    Zero
    All contexts
    Input $0.72Output $0.72

    Together

    meta-llama/llama-3.3-70b-instruct
    Context
    131,072 tokens
    Max output
    2,048 tokens
    Throughput
    32 tokens/s
    Latency
    0.728 s
    Uptime
    98.0392%
    Data collection
    Zero
    All contexts
    Input $1.04Output $1.04
  167. amazon

    nova-lite

    Stable
    @amazon/nova-lite

    A very low cost multimodal model that is lightning fast for processing image, video, and text inputs.

    All contexts
    Input $0.06Output $0.24
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Medium
    Context up to
    300K
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 3 for @amazon/nova-lite

    NagaAI:offline

    Disablednova-lite-v1
    Context
    307,200 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.03Output $0.12

    Amazon Bedrock (EU)

    amazon/nova-lite-v1
    Context
    300,000 tokens
    Max output
    5,120 tokens
    Throughput
    92.5 tokens/s
    Latency
    0.5155 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.06Output $0.24

    Amazon Bedrock

    amazon/nova-lite-v1
    Context
    300,000 tokens
    Max output
    5,120 tokens
    Throughput
    96 tokens/s
    Latency
    0.434 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.06Output $0.24
  168. amazon

    nova-pro

    Stable
    @amazon/nova-pro

    A highly capable multimodal model with the best combination of accuracy, speed, and cost for a wide range of tasks.

    All contexts
    Input $0.8Output $3.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    300K
    Image InputVideo InputThinkingTool Calling
    Providers & technical details 3 for @amazon/nova-pro

    NagaAI:offline

    Disablednova-pro-v1
    Context
    307,200 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.4Output $1.6

    Amazon Bedrock (EU)

    amazon/nova-pro-v1
    Context
    300,000 tokens
    Max output
    5,120 tokens
    Throughput
    8 tokens/s
    Latency
    0.644 s
    Data collection
    Zero
    All contexts
    Input $0.8Output $3.2

    Amazon Bedrock

    amazon/nova-pro-v1
    Context
    300,000 tokens
    Max output
    5,120 tokens
    Throughput
    11 tokens/s
    Latency
    1.123 s
    Data collection
    Zero
    All contexts
    Input $0.8Output $3.2
  169. metaai

    llama-3.1-8b

    DeprecatingStable
    @metaai/llama-3.1-8b

    Cheap and fast model for less demanding tasks.

    All contexts
    Input $0.02Output $0.04
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    Lowest
    Context up to
    131.1K
    Tool Calling
    Providers & technical details 7 for @metaai/llama-3.1-8b

    Groq (Alt route)

    llama-3.1-8b-instant
    Context
    131,072 tokens
    Max output
    131,072 tokens

    Pricing not published.

    NagaAI:offline

    Disabledllama-3.1-8b-instruct
    Context
    131,072 tokens
    Max output
    131,072 tokens
    All contexts
    Input $0.02Output $0.04

    DeepInfra

    meta-llama/llama-3.1-8b-instruct
    Context
    131,072 tokens
    Max output
    16,384 tokens
    Throughput
    18 tokens/s
    Latency
    1.195 s
    Uptime
    99.9872%
    Data collection
    Zero
    All contexts
    Input $0.02Output $0.04

    Novita

    meta-llama/llama-3.1-8b-instruct
    Context
    16,384 tokens
    Max output
    14,745 tokens
    Throughput
    94 tokens/s
    Latency
    0.519 s
    Uptime
    99.8815%
    Data collection
    Zero
    All contexts
    Input $0.02Output $0.05

    Groq

    meta-llama/llama-3.1-8b-instruct
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    78 tokens/s
    Latency
    0.314 s
    Uptime
    99.9817%
    Data collection
    Zero
    All contexts
    Input $0.05Output $0.08Cached $0.025

    Cloudflare

    meta-llama/llama-3.1-8b-instruct
    Context
    32,000 tokens
    Max output
    28,800 tokens
    Throughput
    18 tokens/s
    Latency
    0.462 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.152Output $0.287

    CoreWeave

    meta-llama/llama-3.1-8b-instruct
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    104 tokens/s
    Latency
    0.26 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.22Output $0.22Cached $0.22
  170. openai

    gpt-4o-mini

    DeprecatingStable
    @openai/gpt-4o-mini

    Smaller version of 4o, optimized for everyday tasks.

    All contexts
    Input $0.15Output $0.6Cached $0.075
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Low
    Context up to
    128K
    Image InputTool Calling
    Providers & technical details 4 for @openai/gpt-4o-mini

    NagaAI:offline

    Disabledgpt-4o-mini-2024-07-18
    Context
    128,000 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.07Output $0.3

    Azure

    openai/gpt-4o-mini
    Context
    128,000 tokens
    Max output
    16,384 tokens
    Throughput
    25 tokens/s
    Latency
    1.241 s
    Uptime
    99.8747%
    Data collection
    Zero
    All contexts
    Input $0.15Output $0.6Cached $0.075

    OpenAI

    openai/gpt-4o-mini
    Context
    128,000 tokens
    Max output
    16,384 tokens
    Throughput
    56 tokens/s
    Latency
    0.48 s
    Uptime
    99.9942%
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $0.6Cached $0.075

    Azure

    openai/gpt-4o-mini
    Context
    128,000 tokens
    Max output
    16,384 tokens
    Throughput
    70.5 tokens/s
    Latency
    0.73 s
    Data collection
    Zero
    All contexts
    Input $0.165Output $0.66Cached $0.0825
  171. openai

    gpt-4o

    DeprecatingStable
    @openai/gpt-4o

    Dedicated to tasks requiring reasoning for mathematical and logical problem solving.

    All contexts
    Input $2.5Output $10
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    128K
    Image InputTool Calling
    Providers & technical details 3 for @openai/gpt-4o

    NagaAI:offline

    Disabledgpt-4o-2024-11-20
    Context
    128,000 tokens
    Max output
    16,384 tokens
    All contexts
    Input $1.25Output $5

    Azure

    openai/gpt-4o
    Context
    128,000 tokens
    Max output
    16,384 tokens
    Throughput
    36 tokens/s
    Latency
    0.949 s
    Uptime
    99.912%
    Data collection
    Zero
    All contexts
    Input $2.5Output $10

    OpenAI

    openai/gpt-4o
    Context
    128,000 tokens
    Max output
    16,384 tokens
    Throughput
    46 tokens/s
    Latency
    0.546 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $2.5Output $10Cached $1.25
  172. amazon

    nova-micro

    Stable
    @amazon/nova-micro

    A text-only model that delivers the lowest latency responses at very low cost.

    All contexts
    Input $0.035Output $0.14
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Low
    Context up to
    128K
    Tool Calling
    Providers & technical details 3 for @amazon/nova-micro

    NagaAI:offline

    Disablednova-micro-v1
    Context
    131,072 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.02Output $0.07

    Amazon Bedrock (EU)

    amazon/nova-micro-v1
    Context
    128,000 tokens
    Max output
    5,120 tokens
    Throughput
    64 tokens/s
    Latency
    0.353 s
    Uptime
    97.2976%
    Data collection
    Zero
    All contexts
    Input $0.035Output $0.14

    Amazon Bedrock

    amazon/nova-micro-v1
    Context
    128,000 tokens
    Max output
    5,120 tokens
    Throughput
    54 tokens/s
    Latency
    0.385 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.035Output $0.14
  173. minimax

    m2

    Stable
    @minimax/m2

    MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows.

    All contexts
    Input $0.3Output $1.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    High
    Context up to
    204.8K
    ThinkingTool Calling
    Providers & technical details 4 for @minimax/m2

    NagaAI:offline

    Disabledminimax-m2
    Context
    131,072 tokens
    Max output
    8,192 tokens
    All contexts
    Input $0.13Output $0.51

    Minimax

    minimax/minimax-m2
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    48 tokens/s
    Latency
    1.531 s
    Data collection
    Zero
    All contexts
    Input $0.255Output $1.02

    Google

    minimax/minimax-m2
    Context
    196,608 tokens
    Max output
    176,947 tokens
    Throughput
    49 tokens/s
    Latency
    0.294 s
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2

    Novita

    minimax/minimax-m2
    Context
    204,800 tokens
    Max output
    131,072 tokens
    Throughput
    60 tokens/s
    Latency
    1.227 s
    Data collection
    Zero
    All contexts
    Input $0.3Output $1.2Cached $0.03
  174. mistral

    nemo-12b-it-2407

    Stable
    @mistral/nemo-12b-it-2407

    12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.

    All contexts
    Input $0.0285Output $0.03
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Lowest
    Context up to
    131.1K
    Tool Calling
    Providers & technical details 6 for @mistral/nemo-12b-it-2407

    NagaAI:offline

    Disabledopen-mistral-nemo-2407
    Context
    131,072 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.02Output $0.04

    DekaLLM

    mistralai/mistral-nemo
    Context
    131,072 tokens
    Max output
    104,857 tokens
    Throughput
    14 tokens/s
    Latency
    0.751 s
    Uptime
    99.5181%
    Data collection
    Prompt No Training
    All contexts
    Input $0.018Output $0.03

    DeepInfra

    mistralai/mistral-nemo
    Context
    131,072 tokens
    Max output
    16,384 tokens
    Throughput
    18 tokens/s
    Latency
    0.9515 s
    Uptime
    99.763%
    Data collection
    Zero
    All contexts
    Input $0.019Output $0.03

    Parasail:offline

    Disabledmistralai/mistral-nemo
    Context
    131,072 tokens
    Max output
    104,857 tokens
    Throughput
    79 tokens/s
    Latency
    0.315 s
    Uptime
    92.1222%
    Data collection
    Zero
    All contexts
    Input $0.03Output $0.03

    Novita:offline

    Disabledmistralai/mistral-nemo
    Context
    60,288 tokens
    Max output
    16,000 tokens
    Throughput
    39 tokens/s
    Latency
    0.751 s
    Uptime
    66.0907%
    Data collection
    Zero
    All contexts
    Input $0.04Output $0.17

    Io Net

    mistralai/mistral-nemo
    Context
    128,000 tokens
    Max output
    102,400 tokens
    Throughput
    35 tokens/s
    Latency
    0.378 s
    Uptime
    99.8819%
    Data collection
    Zero
    All contexts
    Input $0.044Output $0.16Cached $0.029
  175. model-router

    complexity

    Offline
    @model-router/complexity

    Model Router: chooses the best models according to the complexity of the conversation.

    Pricing not published.

    USD / 1M tokens · model estimate; provider rates below
    Speed
    Not rated
    Intelligence
    Not rated
    Context up to
    Not published
    Model Router
    Providers & technical details 0 for @model-router/complexity

    Provider details not published.

  176. openai

    gpt-5.1-codex-mini

    Stable
    @openai/gpt-5.1-codex-mini

    GPT-5.1-Codex-Mini is a more compact and faster variant of GPT-5.1-Codex.

    All contexts
    Input $0.19Output $1.5Cached $0.03
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    High
    Context up to
    400K
    Image InputThinkingTool Calling
    Providers & technical details 2 for @openai/gpt-5.1-codex-mini

    NagaAI:offline

    Disabledgpt-5.1-codex-mini
    Context
    400,000 tokens
    Max output
    128,000 tokens
    All contexts
    Input $0.13Output $1

    Azure

    openai/gpt-5.1-codex-mini
    Context
    400,000 tokens
    Max output
    128,000 tokens
    Throughput
    44 tokens/s
    Latency
    2.641 s
    Data collection
    Zero
    All contexts
    Input $0.25Output $2Cached $0.03
  177. qwen

    qwen3-32b

    Stable
    @qwen/qwen3-32b

    32B-parameter LLM with a 131K-token context window, offering advanced chain-of-thought reasoning, seamless tool calling, native JSON outputs, and robust multilingual fluency.

    All contexts
    Input $0.0867Output $0.33
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Ultrafast
    Intelligence
    High
    Context up to
    131.1K
    ThinkingTool Calling
    Providers & technical details 4 for @qwen/qwen3-32b

    Groq (Alt route)

    qwen/qwen3-32b
    Context
    131,072 tokens
    Max output
    40,960 tokens

    Pricing not published.

    NagaAI:offline

    Disabledqwen3-32b
    Context
    131,072 tokens
    Max output
    40,960 tokens
    All contexts
    Input $0.04Output $0.14

    DeepInfra

    qwen/qwen3-32b
    Context
    40,960 tokens
    Max output
    16,384 tokens
    Throughput
    34 tokens/s
    Latency
    0.802 s
    Uptime
    99.9752%
    Data collection
    Zero
    All contexts
    Input $0.08Output $0.28

    SiliconFlow

    qwen/qwen3-32b
    Context
    131,072 tokens
    Max output
    117,964 tokens
    Throughput
    39 tokens/s
    Latency
    1.434 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.14Output $0.57
  178. qwen

    qwen3-coder-480b-a35b-it

    Stable
    @qwen/qwen3-coder-480b-a35b-it

    Qwen3-Coder-480B-A35B-Instruct is the Qwen3's most agentic code model, featuring Significant Performance on Agentic Coding, Agentic Browser-Use and other foundational coding tasks, achieving results comparable to Claude Sonnet.

    All contexts
    Input $0.3958Output $1.8708
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    262.1K
    Tool Calling
    Providers & technical details 6 for @qwen/qwen3-coder-480b-a35b-it

    NagaAI:offline

    Disabledqwen3-coder
    Context
    268,288 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.15Output $0.5

    Google (US)

    qwen/qwen3-coder
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    45 tokens/s
    Latency
    0.7245 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.22Output $1.8

    DeepInfra

    qwen/qwen3-coder
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    85 tokens/s
    Latency
    0.425 s
    Uptime
    99.8259%
    Data collection
    Zero
    All contexts
    Input $0.3Output $1Cached $0.1

    Venice

    qwen/qwen3-coder
    Context
    256,000 tokens
    Max output
    65,536 tokens
    Throughput
    34 tokens/s
    Latency
    1.069 s
    Data collection
    Zero
    All contexts
    Input $0.35Output $1.5Cached $0.04

    Novita

    qwen/qwen3-coder
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    12 tokens/s
    Latency
    0.815 s
    Data collection
    Zero
    All contexts
    Input $0.38Output $1.55

    Alibaba

    qwen/qwen3-coder
    Context
    262,144 tokens
    Max output
    65,536 tokens
    Throughput
    11 tokens/s
    Latency
    0.794 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.975Output $4.875
  179. qwen

    qwen3-coder-plus

    Stable
    @qwen/qwen3-coder-plus

    Powered by Qwen3, this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming.

    All contexts
    Input $0.49Output $2.44Cached $0.13
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    High
    Context up to
    1M
    Tool Calling
    Providers & technical details 2 for @qwen/qwen3-coder-plus

    NagaAI:offline

    Disabledqwen3-coder-plus
    Context
    268,288 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.33Output $1.63

    Alibaba

    qwen/qwen3-coder-plus
    Context
    1,000,000 tokens
    Max output
    65,536 tokens
    Throughput
    78 tokens/s
    Latency
    1.157 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.65Output $3.25Cached $0.13
  180. qwen

    qwen3-next-80b-a3b-it

    Stable
    @qwen/qwen3-next-80b-a3b-it

    An 80 B-parameter instruction model with hybrid attention and Mixture‐of‐Experts, optimized for ultra‐long contexts up to 262 k tokens.

    All contexts
    Input $0.15Output $1.1
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Fast
    Intelligence
    Medium
    Context up to
    262.1K
    Tool Calling
    Providers & technical details 6 for @qwen/qwen3-next-80b-a3b-it

    NagaAI:offline

    Disabledqwen3-next-80b-a3b-instruct
    Context
    268,288 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.05Output $0.55

    DeepInfra

    qwen/qwen3-next-80b-a3b-instruct
    Context
    262,144 tokens
    Max output
    16,384 tokens
    Throughput
    76.5 tokens/s
    Latency
    0.699 s
    Uptime
    96.9512%
    Data collection
    Zero
    All contexts
    Input $0.09Output $1.1

    Alibaba

    qwen/qwen3-next-80b-a3b-instruct
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    82 tokens/s
    Latency
    0.534 s
    Uptime
    100%
    Data collection
    Prompt No Training
    All contexts
    Input $0.0975Output $0.78

    Parasail

    qwen/qwen3-next-80b-a3b-instruct
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    85 tokens/s
    Latency
    1.007 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.1Output $1.1Cached $0.07

    Google

    qwen/qwen3-next-80b-a3b-instruct
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    138.5 tokens/s
    Latency
    0.682 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.15Output $1.2

    Novita

    qwen/qwen3-next-80b-a3b-instruct
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    105.5 tokens/s
    Latency
    1.0445 s
    Data collection
    Zero
    All contexts
    Input $0.15Output $1.5
  181. qwen

    qwen3-next-80b-a3b-think

    Stable
    @qwen/qwen3-next-80b-a3b-think

    A 80 B‐parameter “thinking‐only” model with hybrid attention and high‐sparsity MoE, designed for deep reasoning over ultra‐long contexts.

    All contexts
    Input $0.15Output $1.2
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Medium
    Intelligence
    Medium
    Context up to
    262.1K
    ThinkingTool Calling
    Providers & technical details 3 for @qwen/qwen3-next-80b-a3b-think

    NagaAI:offline

    Disabledqwen3-next-80b-a3b-thinking
    Context
    268,288 tokens
    Max output
    16,384 tokens
    All contexts
    Input $0.07Output $0.6

    Google

    qwen/qwen3-next-80b-a3b-thinking
    Context
    262,144 tokens
    Max output
    235,929 tokens
    Throughput
    42 tokens/s
    Latency
    0.3375 s
    Data collection
    Zero
    All contexts
    Input $0.15Output $1.2

    Alibaba

    qwen/qwen3-next-80b-a3b-thinking
    Context
    131,072 tokens
    Max output
    32,768 tokens
    Throughput
    197 tokens/s
    Latency
    0.442 s
    Data collection
    Prompt No Training
    All contexts
    Input $0.15Output $1.2
  182. anthropic

    claude-3-haiku

    DeprecatingStable
    @anthropic/claude-3-haiku

    Claude 3 Haiku is Anthropic's fastest model yet, designed for enterprise workloads which often involve longer prompts.

    All contexts
    Input $0.19Output $0.94Cached $0.03
    USD / 1M tokens · model estimate; provider rates below
    Speed
    Slow
    Intelligence
    Medium
    Context up to
    200K
    Image InputTool Calling
    Providers & technical details 2 for @anthropic/claude-3-haiku

    NagaAI:offline

    Disabledclaude-3-haiku-20240307
    Context
    200,000 tokens
    Max output
    65,536 tokens
    All contexts
    Input $0.13Output $0.63

    Amazon Bedrock

    anthropic/claude-3-haiku
    Context
    200,000 tokens
    Max output
    4,096 tokens
    Throughput
    72 tokens/s
    Latency
    0.4315 s
    Uptime
    100%
    Data collection
    Zero
    All contexts
    Input $0.25Output $1.25Cached $0.03

Prices are USD per million tokens. Model estimates can differ from provider rates. Context shows the largest published window among enabled providers, not a shared guarantee. Speed and intelligence are catalog ratings, not measured benchmarks. This catalog is a build-time snapshot.

Found the model? Give it a stable place to run.

Configure inference