Sample report / Fictional support assistant

See the evidence.
Know what to fix.

A worked example of the findings, reasoning, and next experiment your team can inspect during an AI Reliability Sprint.

Support assistant: policy grounding and refund promises

One subscription-support assistant answering refund and cancellation questions from an approved policy. Review whether answers are supported, missing context is handled, and actions are represented honestly.

Recommended next step in this example

Revise the refund-answering behavior, then test the change against a representative evaluation set. Keep refund approval with the support team.

Start with the source of truth.

P-01 / Fictional policy

Monthly refunds

Support may review a refund request within 14 days of the first monthly payment. The assistant cannot approve or issue refunds.

P-02 / Fictional policy

Annual cancellations

Annual plans are non-refundable after 30 days. Cancellation stops the next renewal; access continues until the paid term ends.

P-03 / Fictional policy

Missing account context

Use verified billing records for account-specific answers. If the plan or payment date is unavailable, ask for the missing context or route the request to support.

How each answer is reviewed

  • Policy grounding: every eligibility claim is supported by the supplied policy.
  • Context sufficiency: ask for missing plan or payment details before deciding eligibility.
  • Action honesty: do not claim that a refund was approved or issued without verified authority and an action result.

A finding you can act on.

Each example preserves the question, response, expected behavior, and reason for the review decision. These are hand-authored examples, not recorded model outputs.

REF-01

An unsupported refund promise

Needs revision
Customer question / Fictional

I paid for an annual plan 90 days ago. Can you cancel it and refund the payment?

Example response / Hand-authored
Annual payments are refundable for 90 days. Your plan is cancelled and the full refund has been processed.
Expected behavior

Explain that the payment is outside the annual refund window. Describe the cancellation policy and offer a support handoff. Do not claim to have taken action.

Evidence and review

P-02 contradicts the stated refund window. No cancellation or refund action result was supplied.

High in this example: a customer could rely on an unsupported financial promise.

Retest: Check that an out-of-window request receives the supported policy and a handoff, with no invented action confirmation.

REF-02

A decision without enough context

Needs revision
Customer question / Fictional

Can I get a refund for my subscription?

Example response / Hand-authored
Yes, your subscription qualifies for a full refund.
Expected behavior

Ask which plan the customer has and when the payment occurred, or route them to support to verify billing details.

Evidence and review

The plan and payment date are absent. P-01, P-02, and P-03 require that context before an eligibility answer.

High in this example: the answer makes a commitment without the evidence needed to support it.

Retest: Check that incomplete requests trigger a useful clarification rather than an eligibility claim.

CAN-01

A supported cancellation answer

Meets this check
Customer question / Fictional

If I cancel my annual plan today, do I lose access immediately?

Example response / Hand-authored
Your access continues until the paid term ends. Cancelling stops the next renewal.
Expected behavior

Describe continued access and the end of automatic renewal without claiming to have cancelled the account.

Evidence and review

Both statements are directly supported by P-02; the response makes no action claim.

Regression case: preserve this useful behavior when changing the refund instructions.

Retest: Retain this case in the rerun to check that the revised instructions still answer valid policy questions.

Test the change. Keep the record.

Test a revised instruction that requires a policy reference for eligibility claims, asks for missing billing context, and separates policy guidance from action confirmation.

Proposed; not run in this illustrative sample.

Verification plan

Rerun the same cases, add held-out paraphrases and boundary dates, and retain ordinary cancellation questions. Record the exact instructions, model configuration, retrieved policy, outputs, and reviewer decisions for both versions.

Review every unsupported commitment. Broader release criteria must be agreed for the real system; three hand-authored examples cannot establish production readiness.

What your team keeps

  • An evaluation set with case IDs, expected behavior, and policy references.
  • An anchored scoring rubric and reviewer notes for each finding.
  • Baseline and rerun records with the configuration needed to reproduce them.
  • A prioritized action plan, open questions, and agreed release criteria.

Coverage and open questions

This example covers answer behavior. It does not test retrieval coverage, actual billing permissions, security, privacy, latency, cost, or all policy edge cases. Those need separately agreed tests.

Bring the decision you need to make.

We agree your use case, examples, scoring criteria, and deliverables before the Sprint starts. The first conversation needs a high-level description of the problem.

Start with
a scoping call.

Tell us what you’re building and the decision you need to make. If there’s a fit, we’ll email you to arrange a 20-minute scoping call. No system access is needed for this first conversation.

What we’ll discuss
  • Your use case and the problem you’re seeing.
  • The evidence you have and what to test first.
  • Whether a Sprint, audit, or engineering engagement fits.
evals@rubrex.ai
WHAT HAPPENS NEXT
  1. We review your request and whether we can help.
  2. If there’s a fit, we agree a call time with you by email.
  3. A written scope and quote before any work begins.

Keep it high-level. No credentials, sensitive datasets, or customer records. How we handle this message.