Research & field notes

Better questions.
More reliable AI.

Research, methods, and practical decisions for teams shipping AI. Read the evidence. Understand the limits. Put it to work.

40briefings
across 7 topics

Start here / Evaluation design

A practical evaluation plan for LLM applications: define success, build a representative dataset, choose graders, and investigate failures before release.

Read the briefing
A repeatable quality bar
01 / Define the decision
02 / Measure the baseline
03 / Investigate the failures
04 / Test the next change

A better score should lead to a better decision.

The library

Find your next question.

40 articles

Model briefings / 5 articles

Inside the current models.

OpenAI, Claude, Meta, Kimi, and Gemini. Dated prices, benchmark settings, and practical evaluation advice. These are research snapshots, not live rankings.

Model briefings3 min read

Claude Opus 5.5: pricing, metrics, and what to test

Review Claude Opus 5.5 with dated pricing and independent planning metrics. Learn how to test coding quality, reviewer effort, and application permissions.

Research briefing · Sources included

General AI / 10 articles

AI decisions, explained.

Ten practical guides for product teams: choosing a use case, understanding costs, protecting customer data, and measuring whether the work improves.

General AI3 min read

How to choose your first B2B AI use case

Choose an AI use case with measurable value, available evidence, and manageable failure costs. A practical scorecard for B2B product teams.

Research briefing · Sources included
General AI3 min read

RAG vs. fine-tuning: diagnose the problem first

Understand when retrieval, prompting, or fine-tuning addresses an AI failure. Compare evidence access, behavior consistency, update costs, and evaluation metrics.

Research briefing · Sources included
General AI3 min read

AI pilot success metrics that survive a real rollout

Measure an AI pilot with accepted outcomes, net handling time, repeat use, and operational costs. Learn why demo accuracy alone cannot justify a rollout.

Research briefing · Sources included

Research methods / 25 articles

Build a repeatable quality bar.

Evaluation design, RAG, agents, human judgment, and production reliability. Follow the evidence from a useful question to a release decision.

Evaluation design3 min read

How to evaluate an LLM application before release

A practical evaluation plan for LLM applications: define success, build a representative dataset, choose graders, and investigate failures before release.

Research briefing · Sources included
Evaluation design3 min read

How many examples does an LLM evaluation need?

Choose an evaluation sample around coverage, uncertainty, and the decision at stake instead of relying on a universal minimum number of test cases.

Research briefing · Sources included
Evaluation design3 min read

Build an LLM evaluation dataset that reflects real work

Design evaluation datasets with task coverage, provenance, held-out examples, and explicit expected behavior. Separate synthetic coverage from real traffic.

Research briefing · Sources included
Evaluation design3 min read

Build an AI failure taxonomy that leads to fixes

Classify AI failures by observable symptoms and causal hypotheses so evaluation findings translate into focused experiments and clear ownership.

Research briefing · Sources included
RAG systems3 min read

When should a RAG assistant say it does not know?

Evaluate answerability, abstention, and clarification together so a RAG assistant avoids unsupported answers without refusing useful work.

Research briefing · Sources included
Human judgment3 min read

Where should limited human evaluation time go?

Allocate human review across representative traffic, uncertain judgments, and consequential failure modes without confusing targeted sampling with prevalence.

Research briefing · Sources included
Production reliability3 min read

LLM regression testing: turn failures into release gates

Build a layered regression suite for prompts, models, retrieval, and tools. Separate deterministic failures from noisy judgments and preserve reproducible runs.

Research briefing · Sources included
Production reliability3 min read

What should an AI evaluation engagement deliver?

A buyer’s guide to evaluation scope, datasets, calibrated scoring, failure analysis, reproducibility, and a handoff your engineering team can use.

Research briefing · Sources included

The model decision desk

Which model fits the task?

Compare independent benchmark results, exact model settings, and API costs in one place.

Explore the model comparison

How to read these notes

Evidence, with its limits attached.

These are research briefings and practical recommendations, not original experimental results. Sources include recent preprints and maintained technical documentation. Illustrative examples are labeled, and each article records when its sources were checked.

Our editorial approach

Put the method to work

What does reliable mean
for your AI system?

Explore AI evaluation