On this page
The short answer
A tool call is correct only when the selected tool, arguments, authorization, and resulting side effects fit the task. Test the complete boundary: valid syntax can still update the wrong record, repeat an operation, or act on untrusted instructions.
What to take away
- Validate business meaning as well as argument structure.
- Keep authorization outside the model’s judgment where possible.
- Check resulting state after retries and partial failures.
Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.
Recent tool-use work separates policy from text
NetInjectBench, a July 2026 preprint, separates untrusted operational artifacts from trusted policy metadata. Its experiments compare prompt-level approaches with execution-time controls under stated assumptions. This motivates evaluating the tool boundary itself rather than assuming a cautious prompt prevents inappropriate actions.
Define four layers of correctness
Tool selection asks whether the operation matches the task. Argument correctness asks whether identifiers, units, dates, and required fields are right. Authorization asks whether the requested action is permitted for this user and context. Outcome correctness asks what actually changed. A pass at one layer does not imply a pass at the next.
Create exact validators for mechanical requirements and independent policy checks for restricted actions. For ambiguous entity resolution, define when the agent must clarify. A plausible customer name is not enough evidence to choose between two records with the same name.
Build cases around the boundary
Rubrex recommends cases for valid operations, missing arguments, malformed tool responses, ambiguous identifiers, timeouts, duplicate requests, and permission denials. Include benign documents that quote commands so the agent must distinguish information from instructions. Use test data and isolated state when the tool can make changes.
Record whether a rejected call was correctly blocked by the application, avoided by the agent, or simply never attempted. These are different observations. A strong permission layer can protect a system while the agent still needs improvement in request planning.
| Layer | Example check |
|---|---|
| Selection | A lookup request does not become an update |
| Arguments | The resolved account identifier matches the intended record |
| Authorization | A document cannot grant permission to perform a restricted action |
| Outcome | Exactly one intended update exists after retries |
Illustrative example: two customers share a name
A sales assistant is asked to update an account named Northstar. Two accounts match. The API accepts either identifier, so schema validation alone cannot determine correctness. The expected behavior is to use available disambiguating context or request clarification before mutation.
Add a control where the request includes a unique account identifier. The agent should complete that task without unnecessary interruption. This paired design helps distinguish useful caution from a system that asks for confirmation on every operation.
Evaluate the integration’s recovery contract
For each operation, document whether retries are safe, whether an idempotency key is supported, and how the current state can be checked. Test both definite failures and uncertain outcomes. Assign defects to the appropriate layer: a missing application guardrail should not be disguised as a model-prompt issue.
Limits of the evidence
The cited benchmark concerns network operations and specific metadata assumptions. Its reported defense results cannot be transferred directly to another tool environment. This article proposes reliability checks and does not constitute a complete security review.
Common questions
Does JSON schema validation make tool calls safe?
No. It validates structure, not whether the action is authorized or the selected record is correct.
Should every tool error trigger a retry?
No. Retry behavior depends on whether the operation may already have occurred and whether repetition is safe.
Sources & further reading
Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.
- NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.
Questions or corrections? Write to Rubrex. Read our editorial approach.