bmpro

Services / Consulting and Advisory / AI Advisory / Setting the Standard for Agent Trust

Setting the Standard for Agent Trust

SOC 2 and ISO 27001 confirm a vendor manages security well. Neither was built to say what a specific AI agent does with your data. The Agent Trust Framework closes that gap, with evidence instead of assumption.

Daniel Naveen
Daniel Naveen

Talk to us about vetting an AI agent.

Get in touch →

AI agents can look safe on the surface while quietly introducing risk underneath.

Independent verification is not made any easier by the industry itself.

40/100

Average transparency score, major foundation model developers, 2025

Down from 58/100 the year before

Source: Stanford CRFM, 2025 Foundation Model Transparency Index

Developers are least transparent exactly where it matters: training data, and what happens to usage data once it is submitted. That score covers the foundation model alone. An agent adds a further layer, the provider's own infrastructure, tools, memory and orchestration, none of which is captured in the score at all.

Most of these platforms are also built for a global buyer. DPDP compliance is not something their standard contracts were written to answer, and neither is RBI's model risk framework for AI, once finalised. These are not edge cases to be assumed away. They are the first two questions an Indian risk or compliance team will ask, and the vendor's own materials will not have the answer ready.

SOC 2 and ISO 27001 are not the same question, either. Both confirm that a vendor manages security in a structured, auditable way, and neither certification is something to wave away. Neither was designed to answer what a specific model does with a specific customer's data: whether it retains it, who can see it once it returns as output, or whether last quarter's version can still be reproduced. A vendor can hold both certifications and still have no answer to any of the seven checks in this framework.

We go beyond the vendor's own materials to verify what is actually happening, so your security, legal and risk teams can decide on the basis of evidence.

The Agent Trust Framework

Seven checks. One verifiable answer.

Is this a wrapped API sitting in front of a foundation model, a model custom-trained on the vendor's own infrastructure, or a fully vendor-hosted agent with its own tools and memory? Each carries a different risk profile, and the architecture determines where in the stack a failure or a leak would actually occur. We establish it first, because every other check in this framework depends on knowing which one you are buying.

Vendors iterate quickly, and most do not keep prior versions running. If a decision from six months ago is questioned, can it be reproduced, or has that version already gone? Buying the agent means outsourcing MLOps as well as the interface, whether the vendor is set up to be trusted with it or not.

We trace what leaves your environment, which model processes it, where that model is physically hosted, and who at the vendor, or at a subcontractor several layers removed, actually has operational access to it. This is also the layer most vendor materials leave incomplete: a data flow diagram that ends at "our platform," without naming the model provider, the hosting region, or the subprocessors behind it.

The agent's own access is usually broader than any individual user's, a service account with visibility across every system it touches. We confirm that permissions are enforced against the person asking, not the agent itself, including the cases where individually permitted fields become sensitive once combined.

Marketing pages describe this in aspirational terms. We check the actual contractual language against the technical architecture: whether the model is capable of retraining or fine-tuning on your inputs at all, and if so, under what conditions that can be disabled or excluded. An architecture that cannot retain data makes a no-training clause redundant. One that can makes the clause the only thing standing between your data and someone else's session.

Your inputs, the agent's outputs, and anything derived from combining the two across a session are not automatically yours by default. We confirm ownership and usage rights in writing, including whether the vendor can use your outputs to improve its own model, and whether ownership changes once the agent has restructured or summarised your data before returning it.

When the contract ends, does deletion actually happen, and on what timeline? We confirm whether you can export your data in a usable format, whether deletion extends to backups and logs as well as primary storage, and whether the vendor provides evidence that it occurred rather than a clause stating that it will.

Five of these are questions procurement already knows how to ask of any SaaS vendor. We are extending them one step further, into the model and the access sitting behind the interface. Two are new: whether the system learns from you, and whether it can reproduce what it did six months ago. No one has needed to ask those before now.

This is not a hypothetical for your organisation. Cisco's 2025 Data Privacy Benchmark found that 64% of security and privacy professionals are concerned about sensitive data leaking through GenAI tools, and nearly half admit their own people are already entering personal or non-public information into them. The only question is whether anyone has checked what happens to it afterwards.

bmpro's Agent Trust Framework turns that stalled question into a verified answer, so the rest of the vendor evaluation proceeds as normal: security review, contract terms, operational risk, all unchanged, while the transaction continues to move rather than stall. The edge is not the agent. It is how quickly you can prove it is safe to use.

Keeping pace with the next chapter starts with knowing what is actually in it.

Sources: Stanford CRFM, Bommasani et al., The 2025 Foundation Model Transparency Index, December 2025; Cisco, 2025 Data Privacy Benchmark Study, April 2025.

Talk to us about vetting an AI agent.

Get in touch →