Governing AI at Scale in Financial Services
Banks and insurers have proven AI works. The next question is whether they can run thousands of agents and still show a regulator exactly what each one did. A seven control framework and a 90 day plan for APRA, MAS, RBI and BNM expectations on Google Cloud.
Founder & Lead Google Cloud Architect
Most banks and insurers have proven AI works. A PoC or two have made it into production. The problem that needs solving now is whether they can run thousands of agents in production and still show a regulator, a board and a customer exactly what each one did and why.
In a pilot, governance is a spreadsheet and a steering committee. At scale, that breaks.
Teams build faster than risk functions can review, vendors ship new model versions monthly, and agents start taking actions rather than just drafting text. The organisations moving fastest are not the ones with the fewest controls. They are the ones that built an agentic platform, so the business can build at scale on top of it, with the controls the regulators need.
What changes at scale
Four things shift when AI moves from a handful of use cases to an operating capability.
- Identity. Using Gemini, or Claude on your desktop is acting as you. Often with permissions that a regulator would not let you give an autonomous agent. This is the single biggest risk for regulated customers. These agents are taking actions not replying in chat.
- Model sprawl. Different teams pick different models, prompts and vector stores. Without a single inventory, nobody can answer the first question a regulator asks: what AI is running, where, and on whose authority.
- Shadow AI. When the approved path is slow, staff paste customer data into consumer tools. Blocking alone does not fix this; a faster sanctioned option does.
- Visibility. Logging chain of thought to see why an agent took an action, have this reviewed constantly, and show this to a regulator if asked.
What regulators now expect
Regulators in Australia, Singapore, India and Malaysia have stopped treating AI as a future topic. None has created an AI licence; all expect existing risk, resilience and security obligations to be applied to AI with more rigour, and they are telling boards so directly.
| Regulator | Instrument | What it asks for |
|---|---|---|
| APRA (Australia) | Letter to industry on AI, 30 April 2026, building on CPS 220, CPS 230 and CPS 234 | Board literacy to challenge AI risk, an AI inventory with lifecycle ownership, mapping of the AI supply chain including fourth parties, exit plans for critical providers, and identity and access controls for non-human actors such as agents |
| MAS (Singapore) | Consultation on Guidelines on AI Risk Management, 13 November 2025 | Board and senior management oversight, an up to date AI inventory, materiality assessment by impact, complexity and reliance, lifecycle controls scaled to risk, third party AI risk management, and explicit coverage of generative AI and agents; a proposed 12 month transition after issue |
| RBI (India) | FREE-AI committee framework, 13 August 2025, and Draft Guidance on Regulatory Principles for Model Risk Management, June 2026 | A board approved model risk framework covering all models including AI and vendor tools, a full model inventory with risk classification, independent validation and bias testing, specific controls for generative AI, human override and kill switch capability, and full accountability when AI is sourced from third parties; consultation closed 24 July 2026 |
| BNM (Malaysia) | Discussion Paper on Artificial Intelligence in the Malaysian Financial Sector, 5 August 2025, alongside existing technology risk (RMiT) and outsourcing requirements | Board level accountability for AI outcomes, a risk based and proportionate approach, fairness and explainability for decisions that affect customers, human oversight of high stakes decisions, and management of third party AI and cloud providers; exploratory paper, feedback closed 17 October 2025, final policy not yet confirmed |
The common thread is simple: know what you run, know who owns it, prove it is controlled, and be able to switch it off or move it.
A control framework that scales
Governance at scale works when every AI system passes through the same seven controls, and most of them are enforced by the platform rather than by a review meeting.
| Control | What good looks like | Evidence a regulator can see |
|---|---|---|
| 1. Inventory | Every model, prompt, agent and data source registered with an owner, purpose and risk tier before it reaches production | A live register, not a quarterly spreadsheet |
| 2. Risk tiering | Use cases rated on customer impact, autonomy and reliance; controls scale with the tier | Documented tier per use case and the approvals it triggered |
| 3. Identity and access | Agents get their own service identities with least privilege, short lived credentials and no shared keys | Access logs tied to a named agent and a named human owner |
| 4. Data boundaries | Sensitive data stays in region, is classified before it reaches a model, and is never used to train third party models | Data residency settings, DLP policies and contract terms |
| 5. Evaluation | Each release tested against accuracy, bias, safety and prompt injection cases before and after deployment | Versioned evaluation results linked to each release |
| 6. Human oversight | Clear rules for what an agent may do alone, what needs approval, and how a person can stop it | Approval records and a tested kill switch |
| 7. Monitoring and audit | Every prompt, tool call and output logged, with drift and anomaly alerts routed to an owner | Immutable logs retained to policy, and incident records |
Two principles hold this together. First, tier before you control: a staff productivity assistant should not wait behind the same review as a credit decisioning agent. Second, make the control the default: if a team has to remember to turn on logging, some will not.
How this looks on Google Cloud
Most of the seven controls map to platform features institutions already use for other regulated workloads, which is why we build AI governance into our Google Cloud foundations:
- One place for models and agents. Gemini, and any partner models from Model Garden (e.g. Claude), are consumed through Gemini Enterprise Agent Platform (formerly Vertex AI) inside the company’s own Google Cloud org, so model access inherits existing IAM, billing, quotas and org policies instead of living on a separate vendor account. Agent Registry holds the approved agents, tools and skills in one place, which becomes the live AI inventory.
- Perimeter and residency. VPC Service Controls and org policies keep data and model calls inside approved projects and regions.
- Every agent has its own identity. Agent Identity gives each agent a unique cryptographic identity tied to IAM authorisation policies, so every action traces back to a named agent and a named human owner, and access can be revoked like any other workload. Agent Gateway then controls which tools and other agents each one can reach, across MCP and A2A connections.
- Screening inputs and outputs. Model Armor, enforced at Agent Gateway, inspects prompts and responses for prompt injection, jailbreak attempts and sensitive data leakage before they reach a model, a tool or a user.
- Evaluation before and after release. Agent Simulation tests an agent against synthetic users and virtualised tools before it ships, scoring task success and safety across multi step conversations. Agent Evaluation then scores live traffic continuously with multi turn autoraters, and results are versioned against each release, which is the evidence control 5 asks for.
- Logging and monitoring by default. Cloud Audit Logs and Agent Observability trace which agent called which model or tool, with what, and why, and feed Google Security Operations. Agent Anomaly Detection and the Agent Security Dashboard in Security Command Center flag unusual agent behaviour to an owner.
- Contract clarity. Google Cloud does not use customer prompts and outputs on these enterprise services to train its models, which simplifies the third party and data privacy assessment.
The point is not the individual features. It is that governance becomes part of the landing zone, so every new use case starts compliant rather than earning it later.
Where to start: the first 90 days
You do not need a two year programme to get control. Three focused phases will close the gaps regulators are most likely to probe.
- Days 1 to 30: find what is running. Build the AI inventory (Wiz can help), including vendor features switched on inside existing SaaS tools. Assign an accountable owner to each item and apply a simple three level risk tier.
- Days 31 to 60: put controls in the platform. Register agents in Agent Registry, route model and tool access through Agent Gateway, turn on logging and Model Armor screening by default, and give every agent its own Agent Identity. Map your AI supply chain and flag material service providers under CPS 230 or the MAS, RBI and BNM equivalents.
- Days 61 to 90: prove it. Run Agent Simulation and Agent Evaluation on your highest tier use cases, test the kill switch, and take a board pack to the risk committee that shows inventory, tiers, controls and open issues on one page.
The takeaway
Governance is what lets a financial services customer say “yes” to the next hundred AI use cases instead of the next five. Build the controls into the platform once, and every team that follows inherits them.
Aviato Consulting helps banks, insurers and wealth managers across APAC build governed AI foundations on Google Cloud, from landing zone to production agents. If you want to see where your current setup stands against APRA, MAS, RBI and BNM expectations, get in touch.
Want to see how your AI governance stacks up?
Book a governance reviewFounder of Aviato Consulting and former Google Cloud Consulting Lead for APAC, specialising in enterprise cloud architecture, APRA CPS 234 and agentic AI.