Psychology12 min read24 July 2026

Stress-Testing the STAR: What 600 Synthetic Personas and 10 Out-of-Character Answers Reveal About the Future of Psychometric Assessment

There is a difference between a tool that works and a tool that works under pressure.

A synthetic stress test of the STAR Cognitive Profile Assessment produced results that should make the established psychometric industry very uncomfortable.


There is a difference between a tool that works and a tool that works under pressure.

Most personality assessments are tested under ideal conditions. Clean data. Attentive respondents. Environments where people have time to think, space to reflect, and no particular reason to rush. Under those conditions, almost anything will produce a coherent result. The question is not whether a framework performs when everything goes right. The question is what happens when things go wrong.

That is the point of a stress test. You do not break things because you expect them to fail. You break them because you need to know where the failure points are, how close they lie to normal operating conditions, and what the consequences look like when the system encounters the kind of noise and inconsistency that real human beings introduce into every assessment they complete.

The STAR Cognitive Profile Assessment has just been through that process. Six hundred synthetic personas. Multiple test runs. Deliberate injection of off-pole noise, simulating respondents who answer inconsistently, distractedly, or under conditions that distort their natural cognitive defaults. The results are not just encouraging. They are, for anyone paying attention to what psychometric assessment is supposed to do, a statement of intent.


What was tested

The stress test was built around the CA2.4 algorithm, the latest and most refined version of the STAR scoring model. CA2.4 incorporates several mechanisms that distinguish it from simpler scoring approaches: weighted tier logic that gives Tier 1 questions (those with the highest diagnostic value) a 2.25x weighting, shadow veto logic that prevents a single dominant pole from masking genuine secondary orientations, and Momentum Quotient analysis that stabilises archetype assignment across different response patterns.

The test protocol was straightforward. Six hundred synthetic personas were generated, each constructed with a defined STAR type and archetype. These were split into two cohorts:

Each persona was run through the full 33-question CPA. The algorithm’s classification was compared against the known archetype to determine accuracy.


The baseline results

Across multiple test runs, the results were consistent to the point of being almost boring:

Set 1: 99% accuracy (596/600 correct, 4 misclassified). Structured Cohort: 99%. Random Cohort: 100%.

Set 2: 100% accuracy (599/600 correct, 1 misclassified).

Set 3: 100% accuracy (600/600 correct, 0 misclassified).

Sets 4-6: 100% accuracy across all remaining runs, with one set showing 599/600.

The worst result across all runs was 99% accuracy. The best was a perfect classification of all 600 personas.

The Random Cohort, in the first run, actually outperformed the Structured Cohort, achieving 100% accuracy against the Structured Cohort’s 99%. This is an important finding. It suggests that the algorithm handles the natural messiness of real human responses at least as well as it handles clean, engineered profiles. In practice, this means the CPA is not a tool that only works on textbook cases. It works on the kind of inconsistent, contextual, occasionally contradictory response patterns that real people produce.

For context, 99-100% classification accuracy across 600 synthetic personas is not normal in the psychometric industry. Most personality assessments, when subjected to rigorous synthetic testing, show meaningful degradation. The STAR CPA did not.


The noise test: where it gets interesting

The baseline results established that the algorithm works under ideal conditions. The more important question is what happens when respondents answer inconsistently.

This is where the noise sensitivity analysis becomes the most revealing piece of data in the entire test, and it is the result that should genuinely concern the established psychometric industry.

The test protocol introduced “off-pole noise” into the synthetic personas. This means that a persona whose true archetype is, say, the Precise Analyst (Thinker/Realist) would answer a percentage of their 33 questions as though they were a different type entirely, selecting options that map to Adventurer or Socialiser rather than their true Thinker/Realist orientation.

The noise scale runs from 0 (perfectly consistent) to 10 (nearly one-third of all questions answered out of character).

Let that number settle. Ten out of thirty-three questions answered in a way that contradicts the respondent’s true archetype. This is not a realistic level of noise for most people. As Dave Chadderton, the creator of the STAR Framework, noted: “I think most people would answer quite consistently, and I’d consider 10/33 to be very noisy.” These synthetic personas are not reasonable approximations of typical respondents. They are simulated flawed human beings, pushed to an extreme that exceeds what most real-world assessment conditions would produce.

The question was not whether the algorithm would degrade. Any algorithm degrades under sufficient noise. The question was how it degrades, and what the pattern of degradation reveals.


What the noise sensitivity chart shows

The per-archetype noise sensitivity analysis plots accuracy for each of the twelve STAR archetypes as off-pole answers increase from 0 to 10. The results divide into two clear categories.

The robust archetypes

The majority of archetypes maintained near-perfect or perfect classification even at maximum noise. These are profiles whose behavioural signature is so distinctive, so clearly differentiated from adjacent types, that even significant answer contamination cannot obscure them. Their motivational fingerprint is loud enough to hear through the static.

These archetypes represent the backbone of the STAR Framework’s classification power. They are the types that, in real-world deployment, will produce reliable results even when respondents are tired, distracted, or answering under time pressure. They are the archetypes where the algorithm’s confidence is highest, and where the gap between one type and its nearest neighbour is wide enough to absorb considerable noise.

The fragile archetype

At the other end of the spectrum sits the Practical Builder (Adventurer/Realist), which experienced what can only be described as a catastrophic degradation under noise. At zero noise, the Practical Builder is classified with 100% accuracy. At maximum noise (10 off-pole answers), accuracy collapses to approximately 10%. A ninety percentage-point drop.

This is not a flaw in the algorithm. It is a feature of the archetype itself, and understanding why it happens tells you something important about how the STAR Framework models human behaviour.

The Practical Builder sits at the intersection of two orientations that, on the surface, appear contradictory: the Adventurer’s drive for autonomy, novelty, and momentum, and the Realist’s need for security, stability, and proven approaches. This tension is not a weakness of the model. It is a genuine reflection of a psychological profile that is inherently more ambiguous than, say, the Precise Analyst (Thinker/Realist), whose two orientations reinforce rather than compete with each other.

When enough answers shift away from that tension point, the algorithm cannot distinguish the Practical Builder from adjacent archetypes whose profiles overlap at the boundaries. The noise does not break the algorithm. It reveals where the behavioural boundaries between archetypes are thinnest, and where the psychological territory is genuinely contested.

This is exactly what you want a psychometric tool to reveal. A tool that hides its uncertainty, that gives you a confident answer even when the data is ambiguous, is a dangerous tool. A tool that tells you where it is confident and where it is not is a trustworthy one.


Why this matters for real-world deployment

Personality assessments do not operate in laboratories. They operate in the real world, where respondents complete them on their phones during lunch breaks, in offices with interruptions, or at home when they are tired and half-thinking about tomorrow’s meeting. The quality of the data that any assessment produces is only as good as the consistency of the respondent’s answers, and the reality is that human beings are not consistent.

The STAR CPA’s noise sensitivity results suggest that the CA2.4 algorithm is built for this reality. When a real respondent answers three or four questions out of thirty-three in a way that does not quite match their true type, the algorithm barely notices. The classification holds. The archetype is correct. The report is accurate.

Even at extreme noise levels that most respondents would never reach, the algorithm degrades gracefully rather than catastrophically. It loses accuracy, but it does not lose coherence. It does not suddenly reclassify a Thinker as an Adventurer because of a handful of inconsistent answers. It recognises the noise, accounts for it, and produces a result that reflects the underlying signal rather than the surface-level static.

With one exception: the Practical Builder, where the genuine psychological ambiguity of the archetype means that heavy noise can push the classification across a boundary that, in the real world, is already somewhat blurred.


The comparison: STAR versus the established frameworks

Dave Chadderton’s question during the review was direct: “How would STAR compare with MBTI, DiSC, and Insights Discovery?”

The honest answer is that direct comparison is difficult, because the frameworks are built on fundamentally different foundations. But the stress test results make a comparison possible, at least in terms of what the data reveals about algorithmic robustness.

Myers-Briggs (MBTI)

MBTI remains the most widely used personality framework in corporate history. It is also, by some distance, the most criticised. Its binary categories (introvert or extrovert, thinker or feeler) force a false choice on a spectrum that is far more nuanced. Studies have shown that up to 50% of people receive a different result when retaking the test just five weeks later.

Under a synthetic stress test of the kind the STAR CPA has undergone, MBTI’s reliability would be expected to degrade substantially. Its categories are not robust to natural response variation, and its forced-choice format means that small shifts in answers can flip the result entirely. The STAR CPA’s tiered weighting system and shadow veto logic exist precisely to prevent this kind of brittleness.

DiSC

DiSC is more practically oriented and deliberately simplified, trading depth for accessibility. It identifies four primary styles and maps behavioural tendencies in workplace contexts. Useful, but shallow. DiSC tells you what someone tends to do. It does not tell you why they do it, what drives that behaviour at a motivational level, or what it costs them when the tendency is overdone.

The STAR CPA’s seven-pillar theoretical foundation provides a depth of explanation that DiSC does not attempt.

Insights Discovery

Insights Discovery uses colour energies derived from Jungian psychology to classify personality preferences. It is visually appealing and intuitive, which explains its popularity in team-building contexts. But the underlying model is essentially a repackaging of MBTI with colour-coded aesthetics. The same reliability concerns apply.

The structural difference

What the STAR CPA’s stress test results demonstrate is that building a psychometric tool on seven established psychological theories, rather than one or two, produces a classification system that is not just more nuanced but more robust. The tiered weighting system, the shadow veto logic, and the Momentum Quotient analysis are not cosmetic features. They are structural mechanisms that allow the algorithm to maintain accuracy under conditions that would compromise simpler frameworks.


The architecture of robustness

The CA2.4 algorithm’s performance under noise is not accidental. It is the product of deliberate design decisions that prioritise robustness over simplicity.

Weighted tier logic: Of the 33 CPA questions, 11 are designated Tier 1 and receive a 2.25x weighting. These questions map most directly to the core motivational drivers identified by Self-Determination Theory and Regulatory Focus Theory. They are the questions where the signal-to-noise ratio is highest, and giving them additional weight ensures that the most diagnostic responses carry proportionally more influence on the final classification.

Shadow veto logic: This mechanism prevents a single dominant pole from overwhelming the secondary orientation. In simpler models, a strong Adventurer score might mask a genuine Realist secondary. The shadow veto ensures that the algorithm checks for suppressed secondary signals, preserving the archetype distinction even when the primary orientation is dominant.

Momentum Quotient analysis: This stabilises classification by examining the consistency of the response pattern across the seven theoretical blocks. A person who answers Adventurer-consistently across all seven blocks has a different Momentum Quotient than someone whose Adventurer responses cluster in some blocks but not others. This additional layer of analysis is what allows the algorithm to maintain accuracy even when individual answers are noisy.

Stability score: The CPA outputs a stability score that tells the respondent how consistent their profile is across different contexts. This is not just a confidence metric. It is an honesty metric. It tells you, and anyone reviewing your results, whether your answers reflect a stable pattern or a noisy one.

Together, these mechanisms create an algorithm that does not just score answers. It interrogates the quality of the answers themselves and adjusts its classification accordingly.


What this means for the future

Dave Chadderton, after reviewing the stress test results, made a declaration: “I believe STAR + DOTS + DCTA is the future of psychometric training and analysis.”

This is not a casual statement. It is a strategic position backed by data.

The STAR Framework provides the psychological model. DOTS (Data, Opportunity, Togetherness, Stabilise) provides the communication framework that translates STAR insights into actionable strategy. DCTA (Drivers, Custodians, Translators, Aligners) provides the team composition framework that applies STAR insights to organisational dynamics.

Together, they form a system that goes far beyond “what personality type are you?” It answers: how do you think, what motivates you, how do you communicate, what derails you, who should you work with, and what does your team need from you?

The stress test results validate the foundation. The CA2.4 algorithm is robust. The archetypes are real. The classification holds under pressure.

What remains is the work of taking it into organisations, teams, and contexts where the established tools have dominated for decades, often without the kind of algorithmic validation that the STAR CPA has just undergone.


A note on what this is not

This analysis is not a claim that the STAR CPA is perfect. No psychometric tool is. The noise sensitivity analysis makes that explicit: some archetypes are more robust than others, and the algorithm’s confidence varies accordingly.

What it is a claim for is transparency. The STAR CPA does not just give you a result. It tells you how confident it is in that result. It tells you whether your answers were consistent or noisy. It tells you which archetypes are most and least robust. It gives you the map and the legend, not just the destination.

In a psychometric industry that has spent decades selling personality labels without showing its working, that is a genuinely different proposition.


The bottom line

Six hundred synthetic personas. Multiple test runs. 99-100% accuracy across every iteration. Graceful degradation under extreme noise. Differential robustness that reveals genuine psychological boundaries rather than hiding them.

The STAR CPA has been stress-tested. It held.

The question now is not whether the algorithm works. It is how fast the rest of the industry will notice.


David Chadderton is the creator of the STAR Framework and the author of three books: The STAR Framework: Rewriting the Rules of Consumer Engagement (NYC Big Book Award 2025), The STAR Operating System: Decode Mindset, Understand Motivation, Transform Human Behaviour, and Dear Algorithm, It’s Not Me, It’s You. He spent his twenties and thirties as a military aviator and instructor, studying how people make decisions when the stakes are highest. He now applies those principles as a Chief Marketing Officer, bringing behavioural science to performance marketing at scale. He writes about human behaviour, AI, and the psychology of decision-making on The Unoptimised Human.

The STAR Framework

If you enjoyed this essay, you'll find the full argument — and the framework behind it — in the book.