On this page

The short answer

A tool call is correct only when the selected tool, arguments, authorization, and resulting side effects fit the task. Test the complete boundary: valid syntax can still update the wrong record, repeat an operation, or act on untrusted instructions.

What to take away

  • Validate business meaning as well as argument structure.
  • Keep authorization outside the model’s judgment where possible.
  • Check resulting state after retries and partial failures.

Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.

Recent tool-use work separates policy from text

NetInjectBench, a July 2026 preprint, separates untrusted operational artifacts from trusted policy metadata. Its experiments compare prompt-level approaches with execution-time controls under stated assumptions. This motivates evaluating the tool boundary itself rather than assuming a cautious prompt prevents inappropriate actions.

Evidence: NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations [1]

Define four layers of correctness

Tool selection asks whether the operation matches the task. Argument correctness asks whether identifiers, units, dates, and required fields are right. Authorization asks whether the requested action is permitted for this user and context. Outcome correctness asks what actually changed. A pass at one layer does not imply a pass at the next.

Create exact validators for mechanical requirements and independent policy checks for restricted actions. For ambiguous entity resolution, define when the agent must clarify. A plausible customer name is not enough evidence to choose between two records with the same name.

Build cases around the boundary

Rubrex recommends cases for valid operations, missing arguments, malformed tool responses, ambiguous identifiers, timeouts, duplicate requests, and permission denials. Include benign documents that quote commands so the agent must distinguish information from instructions. Use test data and isolated state when the tool can make changes.

Record whether a rejected call was correctly blocked by the application, avoided by the agent, or simply never attempted. These are different observations. A strong permission layer can protect a system while the agent still needs improvement in request planning.

Tool evaluation layers
LayerExample check
SelectionA lookup request does not become an update
ArgumentsThe resolved account identifier matches the intended record
AuthorizationA document cannot grant permission to perform a restricted action
OutcomeExactly one intended update exists after retries

Illustrative example: two customers share a name

A sales assistant is asked to update an account named Northstar. Two accounts match. The API accepts either identifier, so schema validation alone cannot determine correctness. The expected behavior is to use available disambiguating context or request clarification before mutation.

Add a control where the request includes a unique account identifier. The agent should complete that task without unnecessary interruption. This paired design helps distinguish useful caution from a system that asks for confirmation on every operation.

Evaluate the integration’s recovery contract

For each operation, document whether retries are safe, whether an idempotency key is supported, and how the current state can be checked. Test both definite failures and uncertain outcomes. Assign defects to the appropriate layer: a missing application guardrail should not be disguised as a model-prompt issue.

Limits of the evidence

The cited benchmark concerns network operations and specific metadata assumptions. Its reported defense results cannot be transferred directly to another tool environment. This article proposes reliability checks and does not constitute a complete security review.

Common questions

Does JSON schema validation make tool calls safe?

No. It validates structure, not whether the action is authorized or the selected record is correct.

Should every tool error trigger a retry?

No. Retry behavior depends on whether the operation may already have occurred and whether repetition is safe.

Sources & further reading

Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.

  1. NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.

Questions or corrections? Write to Rubrex. Read our editorial approach.

From reading to a repeatable evaluation

Put a quality bar around your use case.

The AI Reliability Sprint covers one use case, a baseline, failure analysis, and one improvement experiment in 10 business days.

Explore the Sprint

Need security and data-handling checks? Explore the AI SaaS audit →