In April 2026, the Australian Prudential Regulatory Authority (APRA) warned the financial services industry that as they seek to realize the productivity and efficiency benefits of AI, governance, the associated governance, risk and compliance management of AI is not keeping pace.
This blog outlines the regulators position and how to address the concerns.
The Regulator’s Position
In April 2026 APRA wrote to regulated banks, insurers, and superannuation trustees after reviewing AI adoption and deployment at large entities. The key takeaway is that whilst AI use is accelerating within these organizations, the associated governance, risk and compliance oversight is lagging.
APRA’s observations highlighted several common themes across the industry:
- Accelerating Adoption: AI is rapidly scaling across diverse business functions, necessitating a parallel evolution in risk management.
- Strengthened Governance & Oversight: Boards must improve AI literacy to oversee strategy and challenge management, moving beyond vendor summaries.
- Cyber & Supplier Risks: Organizations face new attack vectors (like prompt injection) and heightened supply chain dependencies, requiring greater visibility. organizations rely heavily on a small number of AI vendors or lack visibility into upstream dependencies, model changes and fourth-party providers.
- Dynamic Assurance: Static reviews are inadequate; entities must implement continuous monitoring, stronger controls, and clear audit and accountability.
- Enforcement: APRA is prepared to take supervisory action where AI risks are not managed appropriately. APRA expects regulated entities to strengthen controls, including AI inventories, clear accountability, lifecycle governance, staff training, strong cyber controls, fallback processes, supplier oversight and continuous assurance.
Taken together, these observations point to a common challenge. As AI becomes embedded in critical business processes, organizations need stronger governance, clearer accountability, continuous monitoring and better evidence to demonstrate that AI systems remain safe, reliable and operating within approved risk tolerances.
For many organizations, the challenge is no longer whether they can build AI. It is whether they can demonstrate that AI can be trusted in production
Governance, risk and compliance is falling behind
It is not surprising that oversight and assurance functions are falling behind. Governance, risk and compliance in financial services was built for predictable software. Processes like fixed period model validations and tests assume that a system remains stable between reviews.
Generative AI represents a fundamental break from predictable, rules-based machine learning. Because LLMs evolve rapidly via frequent, unannounced model updates, static periodic reviews are obsolete. This creates three critical governance gaps identified by APRA:
- Periodic assurance creates blind spots: Annual snapshots fail to capture rapid output drift, risking inaccuracies in credit, fraud, or claims management.
- Human oversight must be provable: Regulators require auditable logs of human intervention in high-risk decisions, which most firms currently lack. If an AI agent makes a borderline decision, the risk team needs the interaction log, the reasoning, and the review step. Most firms currently lack this telemetry.
- The AI supply chain is opaque: Without comprehensive logs of every model call, managing fourth-party risk from foundation models and training data is effectively impossible. Consistent use and evaluation of models, harnesses and prompts across all applications requires a comprehensive log of activity and continuous oversight.
Solving the Governance Problem
Fixing these gaps requires three components: a tool to capture every AI interaction in detail; a database fast enough to ingest and transform these massive volumes of production data being generated; and an analytical framework to turn that data into usable information for management action, executive oversight and regulatory evidence.
This is what NovoFinity, ClickHouse and Langfuse provide together:
Langfuse: Build, Monitor and Optimize AI Applications at Scale
Langfuse, now part of ClickHouse, is a leading platform for LLM observability, evaluations and prompt management. It helps you understand exactly what an LLM application does in production with the detail that management, boards and regulators expect.
For every interaction, Langfuse creates a structured record called a trace. This includes the input, retrieval steps, model calls, prompt versions, and any evaluation scores. If a decision is later disputed, the risk team can reconstruct the entire process. Integration runs into the background without slowing down your systems or interrupting operation. Deployment is fast – most teams complete the set up within hours.
Continuous evaluation
Langfuse can automatically score production traces based on your criteria, such as factual accuracy or policy compliance. You can see shifts in quality over time and set alerts for when thresholds are breached. This provides the continuous monitoring required and ability to capture drift before it causes serious consequences.
Recorded human review
For high-risk decisions, Langfuse can route flagged traces to human reviewers who can offer feedback on reasoning logic and potential improvement. The final judgment is recorded and linked to the original trace, providing proof of oversight.
Prompt version control
Langfuse tracks every version of a prompt used in production. When a prompt is updated, you can compare performance between the old and new versions, creating the change governance audit trail regulators expect.
ClickHouse: Handling AI Data at Scale
Clickhouse is the leading database for AI capable of powering agentic systems with millisecond queries at petabyte scale. Langfuse is built on ClickHouse because ClickHouse is designed for large-scale data. Standard databases like PostgreSQL often struggle with the volume of AI interactions found in financial services. ClickHouse handles this workload efficiently, using columnar compression to make long-term data retention affordable.
Audit integrity
Langfuse offers a strong audit trail feature as part of their enterprise offering. Any change made to the AI model, prompts, or evaluations are logged. Underneath it is the ClickHouse, that appends data rather than updating it in place, so records cannot be silently changed. It also logs every query, creating a clear chain of custody for audit purposes.
NovoFinity: Turning Signals into Evidence
Novofinity, using their data and regulatory expertise, works with the real time data observed and monitored by Langfuse and stored within Clickhouse, ensuring that organizations have analytics and information at their disposal for management to operationalize and optimize, and boards and regulators to review and challenge.
Engineering signals like traces and token counts are useful, but they aren’t fit for management action, executive oversight or regulatory on their own. Using the real time analytics capability of ClickHouse, NovoFinity translates data into analytics and information that can be used for operational, management and oversight use cases demonstrating whether an AI system is operating within your firm’s business and regulatory expectations.
NovoFinity has extensive global experience helping banks, insurers and other regulated financial institutions strengthen data governance and address complex regulatory obligations. Through data quality frameworks, automated controls, lineage, validation and continuous monitoring, NovoFinity applies the governance disciplines that regulators expect to generate trusted, decision-grade analytics and information. These principles now form the foundation for governing enterprise AI, bringing decades of regulated data governance experience into the next generation of intelligent systems.
What This Partnership Delivers for Oversight and Assurance
- Management accountability: Live dashboards show model performance and review rates instead of static reports.
- Continuous monitoring: Scores are generated for every trace, triggering alerts in seconds.
- Human oversight: Annotation queues link reviewer decisions permanently to AI traces.
- Data leakage prevention: Personally identifiable information is masked at ingestion. Additionally Role Based Access Controls trace access by role which means that every query against sensitive trace data is logged and can be traced back to an individual
- Supplier & fourth-party risk: Models and their underlying data may be updated without notification which may result in inconsistent outputs. Every model call records provider, model version, and data type which triggers alerts to oversight functions.
- Model drift & bias detection: Outcomes change as either the model used changes, the version of the model used changes or the underlying data changes. Our processes log and identify these changes with frameworks that monitor for drift and bias.
- Data retention aligned to required standards: Multi-year interaction data is stored and can be queried without archival or cold-tier delays. Long retention is economically viable even with data volumes experienced within financial services.
- Enterprise level data quality: The speed and frequency of reporting is increasing and there is less time available for remediating data quality issues. Real time analytics using data quality frameworks flags incidents inflight for resolution thereby facilitating improved reporting cycle and analysis times.
- Data Sovereignty: Clickhouse cloud offers data residency on your hyperscaler, keeping data within the region. This is the simplest path to getting started. For more isolation, it can be self-hosted in your VPC. ClickHouse Cloud also offers a Bring Your Own Cloud option. Full on-premises deployment is available for air-gapped requirements.
Next Steps
As AI becomes embedded in your critical business processes, you will need stronger governance, clearer accountability, continuous monitoring and better evidence to demonstrate that AI systems remain safe, reliable and operating within approved risk tolerances, the partnership of NovoFinity, ClickHouse and Langfuse can help.
You do not need to instrument every AI application at once. Start with the one or two applications that carry the highest regulatory exposure. Generally anything customer-facing that touches credit, claims, fraud or advice. You can build the evidence trail there first.
The platform allows for an expeditious implementation process. Once data is flowing, NovoFinity helps define the criteria for quality breaches and human reviews based on your specific regulatory obligations.
This approach moves your governance from reactive to proactive. Rather than scrambling to find evidence when asked, you can simply open a dashboard.
APRA has set clear expectations. The priority now is how quickly your firm can close these governance gaps.
Closing the AI Governance Gap
NovoFinity, ClickHouse and Langfuse are helping firms build APRA-aligned observability. Reach out to discuss how we can assist your organization.
Talk to NovoFinity: [email protected]
Contact ClickHouse and Langfuse: clickhouse.com/contact
About the authors
Albert VenterManaging Partner, NovoFinity
Albert Venter is a Data & AI Governance leader and Managing Partner at NovoFinity, specialising in financial services and regulated industries. He helps organisations turn complex data, AI and regulatory challenges into practical, measurable business value. His expertise spans data and AI governance, strategy, BCBS239, automation, cloud transformation and process optimisation, with experience delivering major initiatives across Australia, Africa and Europe.
Muhammad AliSolutions Architect, ClickHouse · Langfuse GTM Lead, APJ
Muhammad Ali (Ali) is a Solutions Architect at ClickHouse, covering the APAC region, where he works with organizations on real-time analytics, observability, and search infrastructure. He also leads Langfuse’s go-to-market efforts in APJ as the region’s first GTM hire. Previously, Ali spent several years at AWS growing the OpenSearch business across APJ. He’s a regular speaker at events including AWS re:Invent, AWS Summit Sydney, and AI Engineer Melbourne, and is based in Melbourne, Australia.

