Discussing the 5 AI Levels, governance best practices, and common pitfalls that stall adoption between Levels 2 and 4. Drawn from a webinar for 250+ credit union executives hosted by NCU-ISAO.

On July 29, I presented a webinar hosted by NCU-ISAO for more than 250 registered credit union executives. For credit unions, and for financial institutions more generally, AI is shifting from a novelty to a necessity, and executives are looking for solid tactics to advance their AI governance and adoption programs.
Our team has learned a lot of nuance from deploying frontier AI models like ChatGPT, Claude, and Copilot with our credit union customers. As we've discovered gaps and shortcomings from off-the-shelf tools, we have continued to add security controls, productivity features, and hands-on training based on end users' real AI adoption needs.
In the webinar session, I distilled our experience into a set of governance best practices and common pitfalls for credit unions, then closed with a practical blueprint that any CU can begin immediately.
I began with the 5 AI Levels that we use to help our customers understand their progress:
We typically meet our customers at Levels 1 and 2, then aim to build momentum toward Level 4 over the course of a year. Reaching Level 5 depends on the quality of your team, but there is no reason a motivated institution could not get there.
Good governance is what moves an organization from Level 1 to Level 2, and it starts with an AI use policy.
We recommend a human-readable document of 3-5 pages that clearly states what is allowed and explicitly prohibits automated lending decisions (a focus area of the CFPB).
The single biggest differentiator we have seen is distinguishing "Internal" models from "External" models, much like the difference between internal and external email:
The distinction between Internal & External models greatly simplifies everything downstream.
I also caution against two extremes:
When deployed into Microsoft Foundry and properly isolated, an internal ChatGPT, Claude, or Copilot model shares the same data governance as a Word document. If you already have an Azure agreement with Microsoft, no new data arrangements are required
We also suggest allowing file uploads for Internal models but not External ones, and treating Internal models as the default workhorse.
I recommend using a hyperscaler service such as Microsoft Azure, which is where most of our customers land.
Nearly every major model can be deployed into your Azure tenant, establishing a perimeter where no data leaves. When data must go out, for a web search perhaps, we typically help our customers deploy data redaction and rehydration, where you replace sensitive values with placeholders on the way out, then restore them when results return, so sensitive information never leaves your network.
User inputs don't always need to be visible to admins, but it should be accessible during an audit through a dedicated auditor role.
Because examiners have not standardized their requests, comprehensive logging lets you prove usage and slice the data however the moment demands. We recommend capturing the following for each interaction:
For MultiModel, we use a simple scale from 0 to 1, where 0 means no sensitive data and 1 means guaranteed sensitive data.
For example, a loan officer summarizing an application document will score a 1 because the prompt contains applicant information. For Internal models, that is acceptable because the system operates without leaving the Azure tenant.
Risk scoring helps teams tune DLP filters early and later demonstrate to examiners that both detection and remediation are in place.
MIT researchers found that 95% of corporate AI initiatives fail. These failures usually happen between AI Level 2 and AI Level 4, which we call the "Zone of Failure."
The top 5 reasons that AI initiatives stall are:
This is the most common failure. Having a few power users "looks like success," but in reality it is three people doing everyone else's work.
The fix is knowing how your team is progressing: track power users, weekly users, entry-level users, and non-users week to week, then run workshops on real use cases and prompt writing to develop a genuine pipeline of talent.
Because AI supposedly "does everything", many people don't know where to start, feel inept, and put it off.
We have found the best way to combat this problem is to deploy targeted RAG models. By connecting specific data sources (e.g. career progression documentation), users can build workflows on top of concrete starting points.
Closely related to the last pitfall, this leaves entry-level users stuck with toy solutions while paying premium rates. Curated RAG models, combined with automated prompt-improvement tools and a shared prompt library, help move these users toward weekly, productive use.
This one is best captured by the $7,000 Claude bill no one saw coming. Providers do not make spending easy to understand, so I recommend tracking cost by user and by department, and eventually setting budgets so users self-regulate.
A key insight here is that much of the cost comes from using overqualified models, such as Opus instead of Sonnet, for simple tasks.
Well-crafted prompts also allow cheaper models to perform well, delivering quality and savings at the same time.
It is generally easy to build a workflow that works 80% of the time but difficult to make a workflow that works 100% of the time.
Rather than trying to perfect a few automations, I advise automating many tasks to 80% while keeping humans in the loop, because humans catch errors AI cannot.
AI can also be "correct but wrong"; for example, an outdated interest rate based on the training data snapshot is technically accurate to the model but is misleading to users. Hence, a direct member-facing automated AI for loans is risky, and striving for full automation can easily cause problems.
I closed the webinar with a blueprint for Levels 1 and 2.
If you are at Level 1, focus on proper governance:
If you are at Level 2, focus on structured adoption:
With our purpose-built software, much of this comes together in the 2 weeks. With off-the-shelf Copilot and Purview, plan for 90 to 180 days.
In closing, the payoff is worth the effort. When AI adoption starts to click, people get very excited. The board is happy, executives are happy, and the users are happy.
AI adoption is a journey, and with good foundational governance and a deliberate path through each AI Level, it is one any institution can navigate.
Deliver leading AI assistants and implement your AI use policies.