AI SaaS security audits / Rubrex

Test the boundaries.
Trace the data.

Security and data-handling assessments for AI applications, SaaS products, and agents. Find where access, actions, and customer data can go wrong, then give your team evidence it can act on.

AI evaluation and engineering is our foundation. Security assessment brings the same discipline to the boundaries around your product.

The model is one part of the system.

An AI SaaS security audit examines the application, its AI workflows, and the data they touch. We agree which checks apply to your product and report untested areas explicitly.

Prompt injection and agent boundaries

Test whether untrusted prompts, retrieved documents, or tool outputs can change an assistant’s authorized task.

  • Direct and indirect prompt injection
  • Tool permissions and approval boundaries
  • Unintended actions, retries, and side effects

Customer and tenant isolation

Check whether users and agents can access only the records, documents, and actions their roles permit.

  • Authentication and object-level authorization
  • Cross-tenant retrieval, caches, and file access
  • Session handling and privileged operations

Data handling and provider settings

Trace customer data through the application, model providers, logging, analytics, and storage.

  • Collection, retention, and deletion paths
  • Training-use settings and supporting provider terms
  • Third-party sharing, logs, and sensitive-data exposure

Application and API security

Review the application around the model, including the interfaces that accept data and execute actions.

  • Input validation and generated-output handling
  • Secrets, dependencies, uploads, and configuration
  • Rate limits, resource limits, and abuse controls

Retrieval and data quality

Evaluate whether the system retrieves permitted, relevant evidence and preserves meaning when it processes data.

  • Source provenance and stale or conflicting documents
  • Citation support and unsupported answers
  • Extraction accuracy, completeness, and schema meaning

Product behavior under failure

Measure how the AI workflow behaves when evidence, tools, or providers fail, and after a proposed fix.

  • Representative task and regression cases
  • Abstention, escalation, and safe failure behavior
  • Repeatability, task completion, and operating limits

What happens to customer data?

Processing a request, retaining a log, sharing data with a provider, and using data for training are different activities. A useful assessment identifies the claim, the evidence, and the limits.

Questions we can investigate within an agreed scope
QuestionEvidence to examineWhat the report distinguishes
Is customer content used for training?Data paths, provider settings, applicable terms, and available records.Observed configuration, contractual statements, and behavior we cannot inspect.
Can one customer access another’s data?Authorization tests across APIs, retrieval, storage, and caches.The tested roles and paths, including any gaps in access or coverage.
What remains after deletion?Deletion workflows, logs, provider retention, backups, and documented exceptions.Verified deletion behavior versus retention that depends on policy or a third party.
Is extracted or generated data correct?Representative reference cases, field-level checks, and source support.Measured quality on the evaluation set, separate from security and privacy findings.

Findings you can investigate. Fixes you can retest.

Scope and data-flow map

The assessed product version, environments, authorized systems, data paths, and exclusions.

Evidence-backed findings

Prioritized issues with affected components, reproduction evidence, business impact, and recommended fixes.

Evaluation and retest record

Agreed cases, observed behavior, configurations, unresolved questions, and results for fixes retested within scope.

Decision-ready handoff

An executive summary, technical report, and next actions. A separate public summary can be prepared when agreed and supported by evidence.

Inspect the report outline

Agree the scope.
Follow the evidence.

Testing starts with written authorization, target systems, permitted actions, and stop conditions. We use synthetic or approved data and record the limits of the available access.

The scope maps relevant checks to OWASP ASVS and considers OWASP guidance for LLM and agentic risks. The report names the versions and requirements actually used.

Read the assessment method
  1. Define the system and claims

    Map the AI workflow, data flows, permissions, and questions the assessment must answer.

  2. Test and review

    Run agreed checks, examine supporting evidence, and separate observed issues from unverified concerns.

  3. Prioritize and retest

    Explain impact, propose fixes, and record the result of agreed retests against the updated system.

  4. Hand over the decision

    Deliver findings, remaining limitations, and a practical next-action plan.

A claim should lead
to its evidence.

A Rubrex assessment badge must identify a specific reviewed product and link to a public record with scope, dates, evidence summary, and current status. It cannot stand for “zero risk” or “no data use.”

Badge eligibility depends on documented criteria and retest evidence. Any material change to the assessed system requires review of the affected claims.

Read the badge requirements

A few useful answers.

What is an AI SaaS security audit?

It is a scoped assessment of an AI product’s application security, model and agent boundaries, and customer-data handling. Rubrex can also evaluate retrieval, structured data, and output quality against agreed requirements. The report separates security findings from quality results.

Can you verify that customer data is not used for model training?

We can review the application’s data flows, relevant provider settings, contracts, and available technical evidence for a precisely defined training-use claim. A provider statement is labeled as such. We cannot prove a provider’s undisclosed internal behavior from an application test or extend a finding beyond the assessed setup.

Is this part of the AI Reliability Sprint?

Security audits are scoped separately. The Sprint remains a 10-business-day evaluation of one use case. If both are needed, we agree how the evaluation cases, security checks, and deliverables fit together before work begins.

Do you use AI to run the audit?

Automation and AI-assisted analysis can support the agreed work. Findings need reproducible evidence and review; model-generated judgments or scanner output alone are not enough to substantiate an assessment claim. Any tools that process client data must be agreed before use.

What access and authorization do you need?

We agree the target systems, permitted tests, test accounts, timing, stop conditions, and data terms in writing. Depending on scope, useful inputs include architecture, code, provider configuration, and a test environment with synthetic or approved data. Access limitations are recorded in the report.

Does an assessment badge guarantee security or compliance?

No. A badge must refer to a specific, dated assessment and a public verification record explaining its scope and status. It cannot promise zero vulnerabilities, establish legal compliance, or stand in for SOC 2 or ISO certification. Payment buys the assessment, not a passing result.

Inspect the reasoning behind the tests.

What isn’t working
the way it should?

Tell us what you’re building, where it breaks, and what you need next. We’ll reply by email to discuss fit and scope.

evals@rubrex.ai
WHAT HAPPENS NEXT
  1. A short email exchange about the problem.
  2. A technical conversation if there’s a fit.
  3. A written scope and quote before any work begins.

Keep it high-level. No credentials, sensitive datasets, or customer records. How we handle this message.