On this page

The short answer

Kimi K3 is worth testing for complex reasoning and long agent tasks, with independent science and mathematics results available in the saved snapshot. Match the exact reasoning setting and harness to your evaluation. Neither a large context window nor a long-running demonstration proves dependable completion in your product.

What to take away

  • Keep Kimi app, Code, and API comparisons separate.
  • Science and math percentages measure different tasks.
  • Measure verified progress during long runs.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

A dated model snapshot. Published metrics describe the named model, settings, and test. They are not Rubrex experiments or a universal ranking. Prices and availability can change after the research date.

Compare models and check the data date →

Separate Kimi K3 from the surrounding products

Moonshot AI’s Kimi K3 technical article describes a model available through Kimi products and its API. Current Kimi Code documentation lists K3 alongside other model options. A Kimi app mode, a coding agent, and an API model are different systems, even when they use the same underlying family. Record the exact model identifier, reasoning effort, tool harness, and access route when evaluating it. Do not compare an entire agent product with a bare model call without stating the difference.

Evidence: Kimi K3: Open Frontier Intelligence [1]Kimi Code model configuration [2]

Read capacity and cost as dated specifications

The saved catalog lists Kimi K3 with 1,048,576 tokens of context and base rates of $3 per million input tokens and $15 per million output tokens. Those values describe the catalog endpoint checked for this article, not the price of a Kimi subscription or a self-hosted deployment. Our recommendation is to estimate cost from complete traces on your workload. Long tasks may accumulate substantial intermediate output, repeated context, and tool activity even when the final answer is short.

Evidence: OpenRouter public model catalog [3]

API catalog snapshot checked September 25, 2026. USD per million tokens, base input/output rates; cached input, tools, long-context tiers, taxes, and routing may change the bill.
Model or propertyPublished valueHow to interpret it
moonshotai/kimi-k3$3 input / $15 outputOpenRouter catalog base rates
Catalog context1,048,576 tokensEndpoint capacity, not a recall guarantee

Compare science and mathematics separately

Epoch AI’s saved GPQA Diamond result covers difficult science questions, while FrontierMath Tiers 1-3 v2 covers mathematics with Python access. Both rows below use Kimi K3 at max reasoning effort, but the percentages have different denominators and cannot be averaged into a meaningful overall intelligence score. Moonshot’s own benchmark footnotes also describe different agent harnesses across tests. Preserve those conditions when reading provider comparisons rather than treating every reported number as the result of an identical experiment.

Evidence: Epoch AI: GPQA Diamond [4]Epoch AI: FrontierMath Tiers 1–3 v2 [5]Kimi K3: Open Frontier Intelligence [1]

Epoch AI internal evaluations, CC BY 4.0. Snapshot checked September 25, 2026. Standard error is not a confidence interval; scores across different benchmarks are not interchangeable.
BenchmarkExact configurationScoreStandard errorRun date
GPQA DiamondKimi K3 (max)93.12%1.49 percentage points2026-07-16
FrontierMath Tiers 1–3 v2Kimi K3 (max)72.18%2.66 percentage points2026-07-17

Test long tasks for useful persistence

A long-running agent should make verified progress, not simply continue producing activity. Prepare tasks with intermediate checkpoints, tool failures, and a clear stopping condition. Inspect whether Kimi recovers from a missing result, repeats an ineffective action, or claims completion before the environment confirms it. Measure accepted completion, progress per tool call, total cost, and recovery after interruption. For document work, add evidence-location checks and conflicting versions. Keep a separate record of what the harness repaired so model and system contributions remain distinguishable.

Limits of the evidence

Rubrex has not reproduced these runs or verified a self-hosted deployment in this article. The launch post’s weight-release plans are not used as evidence of present download availability or license terms. Prices and independent results are a September 25, 2026 snapshot.

Common questions

Can we use these scores to estimate support accuracy?

No. Evaluate support tasks with your documents and acceptance rubric. The published science and mathematics tests cover different capabilities.

What should we log for a Kimi comparison?

Record the model ID, product or API, reasoning effort, harness version, tool budget, accepted outcome, elapsed time, and all billed usage.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. Kimi K3: Open Frontier Intelligence Moonshot AI · 2026 · Provider technical articleReviewed: launch overview and benchmark harness footnotes. Accessed September 25, 2026.
  2. Kimi Code model configuration Moonshot AI · Living reference · Maintained documentationReviewed: current Kimi Code model options. Accessed September 25, 2026.
  3. OpenRouter public model catalog OpenRouter · Living reference · Maintained documentationReviewed: catalog fields; numerical prices taken from the saved September 25 API snapshot. Accessed September 25, 2026.
  4. Epoch AI: GPQA Diamond Epoch AI · 2026 · Independent benchmarkReviewed: benchmark definition and saved internal evaluation records; CC BY 4.0. Accessed September 25, 2026.
  5. Epoch AI: FrontierMath Tiers 1–3 v2 Epoch AI · 2026 · Independent benchmarkReviewed: benchmark definition and saved internal evaluation records; CC BY 4.0. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint