• About
  • Privacy Poilicy
  • Disclaimer
  • Contact
CoinInsight
  • Home
  • Bitcoin
  • Ethereum
  • Regulation
  • Market
  • Blockchain
  • Ripple
  • Future of Crypto
  • Crypto Mining
No Result
View All Result
  • Home
  • Bitcoin
  • Ethereum
  • Regulation
  • Market
  • Blockchain
  • Ripple
  • Future of Crypto
  • Crypto Mining
No Result
View All Result
CoinInsight
No Result
View All Result
Home Regulation

NIST Is Offering a New AI Evaluation Framework, Not Another Compliance Checklist

Coininsight by Coininsight
August 24, 2026
in Regulation
0
189
SHARES
1.5k
VIEWS
Share on FacebookShare on Twitter


NIST’s new draft framework for evaluating AI systems, TEVV-Athlon, should provide a useful reframe for leaders who rely on checklist-style governance, writes Larry Marks, a GRC adviser and consultant. Demonstrating completion of a process isn’t the same as understanding the evaluation, and those who are tempted to turn the framework itself into the latest checklist are making Goodhart’s Law manifest.

One of the things that caught my attention in NIST’s new TEVV-Athlon framework draft for evaluating AI systems is what it does not do. Namely, the framework does not give organizations a standard list of tests every AI system must pass. Instead, NIST recognizes that AI systems, their uses and their encompassing risks vary too much for a single evaluation methodology to work in every situation. Instead, the framework is designed to be adaptable and customizable to the organization’s objectives and the particular AI system being evaluated.

I have spent much of my career working with cybersecurity and technology risk assessments. One lesson I learned is that an organization can become very good at demonstrating that an assessment was completed without necessarily demonstrating that the right risk was assessed. Once a framework becomes established, there is a natural tendency to standardize it. Questions become checkboxes, evidence collected and management approvals documented. Eventually, completing the process can become almost as important as understanding what the process was intended to tell us.

NIST appears to be trying to avoid that problem in this new draft. TEVV-Athlon starts by asking organizations to articulate their objectives and organize the evaluation around them. Only then does the organization determine what should be measured, how those measurements should be performed and what the resulting evidence means. The four stages move from articulate and organize to define and construct, to apply and measure and finally to synthesize and interrogate.

For compliance and risk professionals, this requires a somewhat different mindset. The objective should not be to demonstrate that an AI system has “passed TEVV.” The more important question is whether the evaluation was designed to tell the organization what it actually needed to know about that AI system in its intended use. That may sound like a small distinction, but, in practice, it is a significant one.

A passing score can create the wrong kind of confidence

One of the more interesting parts of the NIST draft is its discussion of Goodhart’s Law: When a measure becomes a target, it can cease to be a good measure. NIST applies this concept to AI evaluation and warns that optimizing a system to perform well against a particular benchmark may not tell us how well that system will perform in the real world. This should get the attention of compliance and risk professionals.

We are accustomed to measurements. We use risk ratings, control effectiveness scores, key risk indicators and other metrics to help management understand risk. These measurements are very useful. They change complicated information into something that can be compared and acted upon. The danger comes when the score becomes the objective.

The same problem can occur with AI. An organization may establish performance thresholds, conduct testing and conclude that an AI system has successfully met its evaluation criteria. Management then sees a passing result and assumes the risk has been addressed. But what exactly passed?

NIST recommends using multiple complementary evaluation approaches and recognizes the importance of real-world testing rather than relying exclusively on benchmarks. That is an important distinction for compliance. Likewise, the purpose of TEVV should not be to produce a score that makes management comfortable but to produce evidence that helps management understand the AI system, including where uncertainty and residual risk remains.

A good evaluation may therefore produce an uncomfortable answer. It may tell management that the AI performs well under certain conditions but that there is insufficient evidence to reach the same conclusion under others. This may be exactly what management needs to know.

Compliance needs to ask a different question

The practical challenge for compliance and risk professionals is resisting the natural desire to reduce TEVV to a simple question: Did the AI pass? I would ask something different: What did this evaluation actually prove? Instead of simply confirming that testing occurred, I want to understand what the organization was trying to learn, why particular measurements were selected, what assumptions influenced the evaluation and what was outside its scope. 

This is consistent with the structure NIST has proposed. TEVV-Athlon starts with organizational objectives and ends with “synthesize and interrogate,” where the results are interpreted to provide information that can support organizational decisions. For compliance professionals, that last step may ultimately be the most important.

We do not need to become data scientists to challenge an AI evaluation. We do need to understand the reasonable conclusions management can draw from it.

I would want four questions answered before relying on the results:

  • What were we trying to learn?
  • Why did we choose these measurements?
  • What assumptions or limitations affected the results?
  • What didn’t we test?

Those questions are different from asking whether the required testing was completed. They require us to examine the relationship between the evidence and the business decision being made. This is particularly important as AI becomes embedded in third-party products. Organizations may receive evaluation reports from vendors rather than conduct every test themselves. A vendor may be able to show impressive benchmark results. The compliance question remains the same: Do those results provide evidence about the way our organization intends to use the AI?

That is where TEVV can become more than another control requirement. Used properly, it can improve the quality of the information management uses to make decisions about AI.

Don’t standardize the value out of TEVV

After years of working with technology risk assessments, I have learned that completing an assessment and understanding risk are not necessarily the same thing. That is why I find NIST AI 200-2 interesting and useful.

NIST is not offering organizations another universal AI test. TEVV-Athlon provides a structured way to determine what an organization needs to know about an AI system and then to construct an evaluation capable of producing evidence relevant to that objective. Its flexibility is not a weakness that compliance departments need to correct. It is part of the framework’s value.

Inevitably, pressures will mount to standardize the process. Organizations want consistency. Auditors want evidence. Executives want understandable results. Regulators want organizations to demonstrate that appropriate controls exist. All of those expectations are reasonable.

The problem begins when demonstrating completion of the process becomes more important than understanding what the evaluation tells us. NIST’s inclusion of Goodhart’s Law should serve as a useful warning. The moment organizations begin managing AI evaluations primarily to achieve acceptable scores, they risk losing sight of why those measurements were selected in the first place.

Related articles

5 key strategies for compliance leaders

August 23, 2026

What’s new in Astute, August 2026 update (v3.5.9)

August 23, 2026


NIST’s new draft framework for evaluating AI systems, TEVV-Athlon, should provide a useful reframe for leaders who rely on checklist-style governance, writes Larry Marks, a GRC adviser and consultant. Demonstrating completion of a process isn’t the same as understanding the evaluation, and those who are tempted to turn the framework itself into the latest checklist are making Goodhart’s Law manifest.

One of the things that caught my attention in NIST’s new TEVV-Athlon framework draft for evaluating AI systems is what it does not do. Namely, the framework does not give organizations a standard list of tests every AI system must pass. Instead, NIST recognizes that AI systems, their uses and their encompassing risks vary too much for a single evaluation methodology to work in every situation. Instead, the framework is designed to be adaptable and customizable to the organization’s objectives and the particular AI system being evaluated.

I have spent much of my career working with cybersecurity and technology risk assessments. One lesson I learned is that an organization can become very good at demonstrating that an assessment was completed without necessarily demonstrating that the right risk was assessed. Once a framework becomes established, there is a natural tendency to standardize it. Questions become checkboxes, evidence collected and management approvals documented. Eventually, completing the process can become almost as important as understanding what the process was intended to tell us.

NIST appears to be trying to avoid that problem in this new draft. TEVV-Athlon starts by asking organizations to articulate their objectives and organize the evaluation around them. Only then does the organization determine what should be measured, how those measurements should be performed and what the resulting evidence means. The four stages move from articulate and organize to define and construct, to apply and measure and finally to synthesize and interrogate.

For compliance and risk professionals, this requires a somewhat different mindset. The objective should not be to demonstrate that an AI system has “passed TEVV.” The more important question is whether the evaluation was designed to tell the organization what it actually needed to know about that AI system in its intended use. That may sound like a small distinction, but, in practice, it is a significant one.

A passing score can create the wrong kind of confidence

One of the more interesting parts of the NIST draft is its discussion of Goodhart’s Law: When a measure becomes a target, it can cease to be a good measure. NIST applies this concept to AI evaluation and warns that optimizing a system to perform well against a particular benchmark may not tell us how well that system will perform in the real world. This should get the attention of compliance and risk professionals.

We are accustomed to measurements. We use risk ratings, control effectiveness scores, key risk indicators and other metrics to help management understand risk. These measurements are very useful. They change complicated information into something that can be compared and acted upon. The danger comes when the score becomes the objective.

The same problem can occur with AI. An organization may establish performance thresholds, conduct testing and conclude that an AI system has successfully met its evaluation criteria. Management then sees a passing result and assumes the risk has been addressed. But what exactly passed?

NIST recommends using multiple complementary evaluation approaches and recognizes the importance of real-world testing rather than relying exclusively on benchmarks. That is an important distinction for compliance. Likewise, the purpose of TEVV should not be to produce a score that makes management comfortable but to produce evidence that helps management understand the AI system, including where uncertainty and residual risk remains.

A good evaluation may therefore produce an uncomfortable answer. It may tell management that the AI performs well under certain conditions but that there is insufficient evidence to reach the same conclusion under others. This may be exactly what management needs to know.

Compliance needs to ask a different question

The practical challenge for compliance and risk professionals is resisting the natural desire to reduce TEVV to a simple question: Did the AI pass? I would ask something different: What did this evaluation actually prove? Instead of simply confirming that testing occurred, I want to understand what the organization was trying to learn, why particular measurements were selected, what assumptions influenced the evaluation and what was outside its scope. 

This is consistent with the structure NIST has proposed. TEVV-Athlon starts with organizational objectives and ends with “synthesize and interrogate,” where the results are interpreted to provide information that can support organizational decisions. For compliance professionals, that last step may ultimately be the most important.

We do not need to become data scientists to challenge an AI evaluation. We do need to understand the reasonable conclusions management can draw from it.

I would want four questions answered before relying on the results:

  • What were we trying to learn?
  • Why did we choose these measurements?
  • What assumptions or limitations affected the results?
  • What didn’t we test?

Those questions are different from asking whether the required testing was completed. They require us to examine the relationship between the evidence and the business decision being made. This is particularly important as AI becomes embedded in third-party products. Organizations may receive evaluation reports from vendors rather than conduct every test themselves. A vendor may be able to show impressive benchmark results. The compliance question remains the same: Do those results provide evidence about the way our organization intends to use the AI?

That is where TEVV can become more than another control requirement. Used properly, it can improve the quality of the information management uses to make decisions about AI.

Don’t standardize the value out of TEVV

After years of working with technology risk assessments, I have learned that completing an assessment and understanding risk are not necessarily the same thing. That is why I find NIST AI 200-2 interesting and useful.

NIST is not offering organizations another universal AI test. TEVV-Athlon provides a structured way to determine what an organization needs to know about an AI system and then to construct an evaluation capable of producing evidence relevant to that objective. Its flexibility is not a weakness that compliance departments need to correct. It is part of the framework’s value.

Inevitably, pressures will mount to standardize the process. Organizations want consistency. Auditors want evidence. Executives want understandable results. Regulators want organizations to demonstrate that appropriate controls exist. All of those expectations are reasonable.

The problem begins when demonstrating completion of the process becomes more important than understanding what the evaluation tells us. NIST’s inclusion of Goodhart’s Law should serve as a useful warning. The moment organizations begin managing AI evaluations primarily to achieve acceptable scores, they risk losing sight of why those measurements were selected in the first place.

Share76Tweet47

Related Posts

5 key strategies for compliance leaders

by Coininsight
August 23, 2026
0

As organizations begin planning for 2027, compliance leaders are balancing growing expectations with limited resources. AI governance, evolving regulations, and...

What’s new in Astute, August 2026 update (v3.5.9)

by Coininsight
August 23, 2026
0

This month’s release focuses on Instructor Led Training. We’ve improved how completed classes appear to learners, made class and course...

Harassment and Inclusion Training: Better Together

by Coininsight
August 22, 2026
0

One HR leader shared a situation with me that I’ve thought about ever since.  A manager wasn’t generating complaints. He...

GRC News Roundup: Speeki, Napier AI, Descartes & More

by Coininsight
August 21, 2026
0

GRC technology is one of the fastest-growing segments in enterprise software, and compliance professions are rapidly evolving. Here’s the latest...

Mitigating tribunal risk: Your questions answered

by Coininsight
August 21, 2026
0

We had so many excellent questions during our recent webinar on the Employment Rights Act and tribunal risks that we...

Load More
  • Trending
  • Comments
  • Latest
What’s Actually Going On With Ripple’s Blockchain?

What’s Actually Going On With Ripple’s Blockchain?

January 12, 2026
MetaMask Launches An NFT Reward Program – Right here’s Extra Data..

MetaMask Launches An NFT Reward Program – Right here’s Extra Data..

July 24, 2025
Finest Bitaxe Gamma 601 Overclock Settings & Tuning Information

Finest Bitaxe Gamma 601 Overclock Settings & Tuning Information

November 26, 2025
Easy methods to Host a Storj Node – Setup, Earnings & Experiences

Easy methods to Host a Storj Node – Setup, Earnings & Experiences

March 11, 2025
Kuwait bans Bitcoin mining over power issues and authorized violations

Kuwait bans Bitcoin mining over power issues and authorized violations

2
The Ethereum Basis’s Imaginative and prescient | Ethereum Basis Weblog

The Ethereum Basis’s Imaginative and prescient | Ethereum Basis Weblog

2
Unchained Launches Multi-Million Greenback Bitcoin Legacy Mission

Unchained Launches Multi-Million Greenback Bitcoin Legacy Mission

1
Earnings Preview: Microsoft anticipated to report larger Q3 income, revenue

Earnings Preview: Microsoft anticipated to report larger Q3 income, revenue

1

NIST Is Offering a New AI Evaluation Framework, Not Another Compliance Checklist

August 24, 2026

Crypto Firm Shut Down After Nine Investors Lose Over £300,000

August 24, 2026

TJX Companies (TJX) Beat Expectations, but the Stock Reaction Shows Investors May Want More Than Defensive Retail Resilience

August 24, 2026

PLTR Price Prediction: Momentum Has Stalled at $180 — A Flush to $172 Is the High-Probability Play Before Bulls Reload

August 24, 2026

CoinInight

Welcome to CoinInsight.co.uk – your trusted source for all things cryptocurrency! We are passionate about educating and informing our audience on the rapidly evolving world of digital assets, blockchain technology, and the future of finance.

Categories

  • Bitcoin
  • Blockchain
  • Crypto Mining
  • Ethereum
  • Future of Crypto
  • Market
  • Regulation
  • Ripple

Recent News

NIST Is Offering a New AI Evaluation Framework, Not Another Compliance Checklist

August 24, 2026

Crypto Firm Shut Down After Nine Investors Lose Over £300,000

August 24, 2026
  • About
  • Privacy Poilicy
  • Disclaimer
  • Contact

© 2025- https://coininsight.co.uk/ - All Rights Reserved

No Result
View All Result
  • Home
  • Bitcoin
  • Ethereum
  • Regulation
  • Market
  • Blockchain
  • Ripple
  • Future of Crypto
  • Crypto Mining

© 2025- https://coininsight.co.uk/ - All Rights Reserved

Social Media Auto Publish Powered By : XYZScripts.com
Verified by MonsterInsights