Security and data-handling assessments for AI applications, SaaS products, and agents. Find where access, actions, and customer data can go wrong, then give your team evidence it can act on.
AI evaluation and engineering is our foundation. Security assessment brings the same discipline to the boundaries around your product.
01The assessment
The model is one part of the system.
An AI SaaS security audit examines the application, its AI workflows, and the data they touch. We agree which checks apply to your product and report untested areas explicitly.
Prompt injection and agent boundaries
Test whether untrusted prompts, retrieved documents, or tool outputs can change an assistant’s authorized task.
Direct and indirect prompt injection
Tool permissions and approval boundaries
Unintended actions, retries, and side effects
Customer and tenant isolation
Check whether users and agents can access only the records, documents, and actions their roles permit.
Authentication and object-level authorization
Cross-tenant retrieval, caches, and file access
Session handling and privileged operations
Data handling and provider settings
Trace customer data through the application, model providers, logging, analytics, and storage.
Collection, retention, and deletion paths
Training-use settings and supporting provider terms
Third-party sharing, logs, and sensitive-data exposure
Application and API security
Review the application around the model, including the interfaces that accept data and execute actions.
Input validation and generated-output handling
Secrets, dependencies, uploads, and configuration
Rate limits, resource limits, and abuse controls
Retrieval and data quality
Evaluate whether the system retrieves permitted, relevant evidence and preserves meaning when it processes data.
Source provenance and stale or conflicting documents
Citation support and unsupported answers
Extraction accuracy, completeness, and schema meaning
Product behavior under failure
Measure how the AI workflow behaves when evidence, tools, or providers fail, and after a proposed fix.
Representative task and regression cases
Abstention, escalation, and safe failure behavior
Repeatability, task completion, and operating limits
02Data claims, made specific
What happens to customer data?
Processing a request, retaining a log, sharing data with a provider, and using data for training are different activities. A useful assessment identifies the claim, the evidence, and the limits.
Questions we can investigate within an agreed scope
Question
Evidence to examine
What the report distinguishes
Is customer content used for training?
Data paths, provider settings, applicable terms, and available records.
Observed configuration, contractual statements, and behavior we cannot inspect.
Can one customer access another’s data?
Authorization tests across APIs, retrieval, storage, and caches.
The tested roles and paths, including any gaps in access or coverage.
What remains after deletion?
Deletion workflows, logs, provider retention, backups, and documented exceptions.
Verified deletion behavior versus retention that depends on policy or a third party.
Is extracted or generated data correct?
Representative reference cases, field-level checks, and source support.
Measured quality on the evaluation set, separate from security and privacy findings.
03What your team keeps
Findings you can investigate. Fixes you can retest.
Scope and data-flow map
The assessed product version, environments, authorized systems, data paths, and exclusions.
Evidence-backed findings
Prioritized issues with affected components, reproduction evidence, business impact, and recommended fixes.
Evaluation and retest record
Agreed cases, observed behavior, configurations, unresolved questions, and results for fixes retested within scope.
Decision-ready handoff
An executive summary, technical report, and next actions. A separate public summary can be prepared when agreed and supported by evidence.
Testing starts with written authorization, target systems, permitted actions, and stop conditions. We use synthetic or approved data and record the limits of the available access.
The scope maps relevant checks to OWASP ASVS and considers OWASP guidance for LLM and agentic risks. The report names the versions and requirements actually used.
Map the AI workflow, data flows, permissions, and questions the assessment must answer.
Test and review
Run agreed checks, examine supporting evidence, and separate observed issues from unverified concerns.
Prioritize and retest
Explain impact, propose fixes, and record the result of agreed retests against the updated system.
Hand over the decision
Deliver findings, remaining limitations, and a practical next-action plan.
05Assessment badges
A claim should lead to its evidence.
A Rubrex assessment badge must identify a specific reviewed product and link to a public record with scope, dates, evidence summary, and current status. It cannot stand for “zero risk” or “no data use.”
Badge eligibility depends on documented criteria and retest evidence. Any material change to the assessed system requires review of the affected claims.
It is a scoped assessment of an AI product’s application security, model and agent boundaries, and customer-data handling. Rubrex can also evaluate retrieval, structured data, and output quality against agreed requirements. The report separates security findings from quality results.
Can you verify that customer data is not used for model training?+
We can review the application’s data flows, relevant provider settings, contracts, and available technical evidence for a precisely defined training-use claim. A provider statement is labeled as such. We cannot prove a provider’s undisclosed internal behavior from an application test or extend a finding beyond the assessed setup.
Is this part of the AI Reliability Sprint?+
Security audits are scoped separately. The Sprint remains a 10-business-day evaluation of one use case. If both are needed, we agree how the evaluation cases, security checks, and deliverables fit together before work begins.
Do you use AI to run the audit?+
Automation and AI-assisted analysis can support the agreed work. Findings need reproducible evidence and review; model-generated judgments or scanner output alone are not enough to substantiate an assessment claim. Any tools that process client data must be agreed before use.
What access and authorization do you need?+
We agree the target systems, permitted tests, test accounts, timing, stop conditions, and data terms in writing. Depending on scope, useful inputs include architecture, code, provider configuration, and a test environment with synthetic or approved data. Access limitations are recorded in the report.
Does an assessment badge guarantee security or compliance?+
No. A badge must refer to a specific, dated assessment and a public verification record explaining its scope and status. It cannot promise zero vulnerabilities, establish legal compliance, or stand in for SOC 2 or ISO certification. Payment buys the assessment, not a passing result.