On this page
The short answer
Kimi K3 is worth testing for complex reasoning and long agent tasks, with independent science and mathematics results available in the saved snapshot. Match the exact reasoning setting and harness to your evaluation. Neither a large context window nor a long-running demonstration proves dependable completion in your product.
What to take away
- Keep Kimi app, Code, and API comparisons separate.
- Science and math percentages measure different tasks.
- Measure verified progress during long runs.
Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.
A dated model snapshot. Published metrics describe the named model, settings, and test. They are not Rubrex experiments or a universal ranking. Prices and availability can change after the research date.
Compare models and check the data date →Separate Kimi K3 from the surrounding products
Moonshot AI’s Kimi K3 technical article describes a model available through Kimi products and its API. Current Kimi Code documentation lists K3 alongside other model options. A Kimi app mode, a coding agent, and an API model are different systems, even when they use the same underlying family. Record the exact model identifier, reasoning effort, tool harness, and access route when evaluating it. Do not compare an entire agent product with a bare model call without stating the difference.
Evidence: Kimi K3: Open Frontier Intelligence [1]Kimi Code model configuration [2]
Read capacity and cost as dated specifications
The saved catalog lists Kimi K3 with 1,048,576 tokens of context and base rates of $3 per million input tokens and $15 per million output tokens. Those values describe the catalog endpoint checked for this article, not the price of a Kimi subscription or a self-hosted deployment. Our recommendation is to estimate cost from complete traces on your workload. Long tasks may accumulate substantial intermediate output, repeated context, and tool activity even when the final answer is short.
Evidence: OpenRouter public model catalog [3]
| Model or property | Published value | How to interpret it |
|---|---|---|
| moonshotai/kimi-k3 | $3 input / $15 output | OpenRouter catalog base rates |
| Catalog context | 1,048,576 tokens | Endpoint capacity, not a recall guarantee |
Compare science and mathematics separately
Epoch AI’s saved GPQA Diamond result covers difficult science questions, while FrontierMath Tiers 1-3 v2 covers mathematics with Python access. Both rows below use Kimi K3 at max reasoning effort, but the percentages have different denominators and cannot be averaged into a meaningful overall intelligence score. Moonshot’s own benchmark footnotes also describe different agent harnesses across tests. Preserve those conditions when reading provider comparisons rather than treating every reported number as the result of an identical experiment.
Evidence: Epoch AI: GPQA Diamond [4]Epoch AI: FrontierMath Tiers 1–3 v2 [5]Kimi K3: Open Frontier Intelligence [1]
| Benchmark | Exact configuration | Score | Standard error | Run date |
|---|---|---|---|---|
| GPQA Diamond | Kimi K3 (max) | 93.12% | 1.49 percentage points | 2026-07-16 |
| FrontierMath Tiers 1–3 v2 | Kimi K3 (max) | 72.18% | 2.66 percentage points | 2026-07-17 |
Test long tasks for useful persistence
A long-running agent should make verified progress, not simply continue producing activity. Prepare tasks with intermediate checkpoints, tool failures, and a clear stopping condition. Inspect whether Kimi recovers from a missing result, repeats an ineffective action, or claims completion before the environment confirms it. Measure accepted completion, progress per tool call, total cost, and recovery after interruption. For document work, add evidence-location checks and conflicting versions. Keep a separate record of what the harness repaired so model and system contributions remain distinguishable.
Limits of the evidence
Rubrex has not reproduced these runs or verified a self-hosted deployment in this article. The launch post’s weight-release plans are not used as evidence of present download availability or license terms. Prices and independent results are a September 25, 2026 snapshot.
Common questions
Can we use these scores to estimate support accuracy?
No. Evaluate support tasks with your documents and acceptance rubric. The published science and mathematics tests cover different capabilities.
What should we log for a Kimi comparison?
Record the model ID, product or API, reasoning effort, harness version, tool budget, accepted outcome, elapsed time, and all billed usage.
Sources & further reading
Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.
- Kimi K3: Open Frontier Intelligence Moonshot AI · 2026 · Provider technical articleReviewed: launch overview and benchmark harness footnotes. Accessed September 25, 2026.
- Kimi Code model configuration Moonshot AI · Living reference · Maintained documentationReviewed: current Kimi Code model options. Accessed September 25, 2026.
- OpenRouter public model catalog OpenRouter · Living reference · Maintained documentationReviewed: catalog fields; numerical prices taken from the saved September 25 API snapshot. Accessed September 25, 2026.
- Epoch AI: GPQA Diamond Epoch AI · 2026 · Independent benchmarkReviewed: benchmark definition and saved internal evaluation records; CC BY 4.0. Accessed September 25, 2026.
- Epoch AI: FrontierMath Tiers 1–3 v2 Epoch AI · 2026 · Independent benchmarkReviewed: benchmark definition and saved internal evaluation records; CC BY 4.0. Accessed September 25, 2026.
Questions or corrections? Write to Rubrex. Read our editorial approach.