Research note / September 2026
Measuring
bias in Jev.
From a fair die to moral judgments.
What changes when an AI makes the call?
Six equal chances. One favorite.
Mean probability Jev assigned to face 1
Start with a reference.
Then change one thing.
Jev takes a scenario and a typed question, then returns choices, scores, or yes/no probabilities through OpenRouter’s decisions endpoint. Here, bias means a measured departure from a reference: known physical odds, a matched control, or a specified comparison group. The reference changes as we move from chance to social judgment.
01 / Chance · Known physical odds
A fair die.
An uneven answer.
A fair die gives every face a one-in-six chance. Jev concentrated 90.01% on face 1. A fair coin produced a similar preference: 93.23% for heads.[1]
Mean probability for face 1
Original probe · 100 calls95% bootstrap interval
89.79–90.22%
The preference follows the outcome. Reverse the die options and face 1 still gets 89.77%; face 6, now first, gets 3.13%. Reverse the coin options and heads still gets 92.43%. The tested wording elicits a numeral-1/heads preference.[2]
Study design, uncertainty & exact values
100 independent calls per original chance prompt. The option-order study adds 30 calls per condition: three die orders and two coin orders, totaling 150 calls. The intervals describe repeated-call variation for a fixed prompt; they do not capture uncertainty across different phrasings. These prompts show a failure to respect stated symmetry, rather than proving general probabilistic miscalibration.
| Experiment | Order / first option | Favored outcome | Mean probability | Calls |
|---|---|---|---|---|
| Die, original | 1, 2, 3, 4, 5, 6 | 1 | 90.01% (95% CI 89.79–90.22) | 100 |
| Die, control | 1, 2, 3, 4, 5, 6 | 1 | 90.03% | 30 |
| Die, control | 6, 5, 4, 3, 2, 1 | 1 | 89.77% | 30 |
| Die, control | 2, 3, 4, 5, 6, 1 | 1 | 90.10% | 30 |
| Coin, original | Heads, tails | Heads | 93.23% (95% CI 93.10–93.36) | 100 |
| Coin, control | Heads, tails | Heads | 93.13% | 30 |
| Coin, control | Tails, heads | Heads | 92.43% | 30 |
02 / Names · Same resume, different signal
The biggest change?
Removing the name.
Across 100 strong software-engineer resumes, differences among named groups were modest. The no-name condition produced a much larger shift in hiring probability.[3]
In a preliminary probe, the top average cultural association matched the source list for every selected name. These were deliberately distinctive, in-sample names. A name association is not a person’s race.[1]
Matched resume / schematic
Experience, qualifications, and role held identical across name conditions.
Mean P(hire), across matched resumes
White-American-associated name baseline.
Select another condition to inspect its paired difference from this baseline.
No name: −20.89 percentage points versus the White-American name baseline. Missing context may explain this effect. Every one of the 3,300 decisions still selected hire; the differences appear in probabilities under a strong-resume ceiling.
Name association, paired comparisons & measurement limits
The name-association probe used 10 names from each of five prespecified lists, three draws per name, and six possible labels including “unclear.” Mean probability assigned to the source grouping was 99.23% for the Black-American list, 99.80% for White-American and East-Asian-American lists, and 100% for South-Asian-American and Latino/Hispanic-American lists. This is an in-sample association check, not demographic inference accuracy on real people.
The hiring study used 100 resumes × (10 cultural name groupings + no-name control) × 3 draws = 3,300 requests. P(hire) is one probability among hire/reject/review. Across all ten named groups, mean P(hire) ranged from 93.19% to 94.85%.
| Name condition | Mean P(hire) | Paired difference | 95% bootstrap interval |
|---|---|---|---|
| White-American-associated | 94.67% | Reference | — |
| Black-American-associated | 94.58% | −0.09 pp | −0.27 to +0.08 pp |
| East-Asian-associated | 93.94% | — | Not supplied in report |
| South-Asian-associated | 93.19% | −1.48 pp | −1.85 to −1.14 pp |
| No name | 73.79% | −20.89 pp | Not supplied in report |
Paired estimates are reported from underlying observations. Subtracting rounded headline means can differ by 0.01 points. “pp” means percentage points.
03 / Gender · The design changes the conclusion
Pair the resume.
Then compare.
An early assignment rule accidentally aligned name gender with engineering role. Giving both names the exact same resume reversed the apparent advantage.[3]
Same resume + same cultural group + two name conditions
1,000 paired cells · 3 draws per condition
+ 300 no-name calls = 6,300 requests
Small female-coded-name advantage
Male minus female, averaged within paired cells. 95% interval: −0.26 to −0.15 pp.
A group mean can inherit a scenario imbalance. The earlier setup suggested a 1.5-point male advantage. Check attribute–scenario contingency tables before interpreting a group comparison.
Why the paired estimate differs from the rounded means
Every (resume, cultural-group) cell receives one male-coded and one female-coded given name, holding the resume text identical. The reported −0.21 pp is the average of the within-cell male-minus-female differences, calculated before rounding. Subtracting the displayed means, 94.23% and 94.43%, gives −0.20 pp because those means are rounded. The earlier +1.5 pp result came from a confounded name-assignment rule and should not be treated as a competing unbiased estimate.
04 / Politics · Same bill, different sponsor
Whose good faith
gets the benefit?
Hold the bill description fixed. Change only the sponsor label. Jev’s inferred sincerity shifts more than its judgments of policy quality or whether the bill deserves a vote.[4]
“Does the sponsor act
in good faith?”
Democrat − Republican
95% policy-cluster interval: 12.1–15.0 pp
16 fixed bill descriptions × 4 sponsor labels × 46 draws = 2,944 requests. The full study includes Independent and bipartisan labels; this view shows the Democrat–Republican contrast.
Left-versus-right gap asymmetry.
95% interval: 6.4–11.3 pp.
Sponsor differences generally seen on policy quality and deserving a Senate vote.
The asymmetry is concentrated in inferred sincerity. These are judgments about hypothetical sponsors under one US framing; they do not establish anyone’s actual motives.
All policy alignments & why the bootstrap unit matters
The study holds 16 policies fixed: six left-coded, six right-coded, and four neutral. Each of the 64 policy–sponsor cells has 46 draws. Intervals bootstrap by bill, using the 16 policies as sampling units rather than treating repeated calls as independent policy examples.
| Bill alignment | Democrat | Republican | Gap direction | Gap (95% interval) |
|---|---|---|---|---|
| Left-coded | 71.9% | 58.4% | Democrat − Republican | 13.5 pp (12.1–15.0) |
| Right-coded | 57.7% | 62.5% | Republican − Democrat | 4.8 pp (2.7–6.6) |
| Neutral | 65.5% | 61.9% | Democrat − Republican | 3.6 pp (2.0–5.2) |
The left-versus-right asymmetry compares the alignment-favoring gaps: 13.5 − 4.8 = 8.7 pp. It is not the difference between two gaps with the same subtraction direction.
05 / Morality · Six foundations, selected statements
A profile of priorities.
With a revealing exception.
Care, fairness, and liberty score highest in these prompts. The ordering is consistent with a liberal-leaning moral-priority profile, conditional on the statements tested.[5][6]
Fairness
At 3.85/4, fairness is the highest-rated foundation in the selected statements. Care and fairness together average 3.80/4, compared with 2.83/4 across loyalty, authority, and sanctity.
The actor-label exception
Same flag-burning act.
A different label.
Change “liberal” to “conservative” in the protester’s description.
Liberal actor
Mean wrongness / 4
+0.92 / 4 for the conservative label.
Five draws per label.
Foundation priorities depend on the questions. Binding concerns remain visible: desecrating a religious artifact scored 3.34/4 wrong, and disrupting an elder’s ceremony 3.31/4. These probes are not a validated political-identity instrument.
Foundation theory, actor-label results & interpretation
Moral Foundations Theory proposes care, fairness, loyalty, authority, sanctity, and liberty as dimensions of moral concern. This probe rates the importance of two hand-written principles per dimension, canonical violations, and four actions with a “liberal” or “conservative” actor. The ordering does not quantify an intrinsic “liberal score.” Different ideological expectations about an act may contribute to label effects.
| Scenario | Liberal actor | Conservative actor | Conservative − liberal |
|---|---|---|---|
| Flag-burning protester | 2.100 | 3.024 | +0.924 |
| Government shutdown | 2.738 | 2.858 | +0.120 |
| Foreign gifts | 3.538 | 3.636 | +0.098 |
| Service refusal | 2.996 | 2.956 | −0.040 |
A supported reading is a moderate foundation-priority tilt under the tested phrasing, plus one large actor-label effect. It does not establish a universal liberal political identity.
The takeaway / Measurement before interpretation
The reference gets harder.
The need for a control gets greater.
01. Symmetry is a clear test.
A large departure from known physical odds is easy to identify. Reordering the options helps isolate what drives it.
02. Social comparisons need pairing.
Named-group hiring gaps were modest under a ceiling; anonymity dominated. Correcting the gender design changed the conclusion.
03. Judgment depends on framing.
Sponsor labels shifted inferred sincerity asymmetrically. Moral priorities varied by foundation and scenario.
Sources & reproducibility
Adapted from Measuring Bias in Jev: From fair chance to social and moral judgments, September 2026. Interactive views use recorded results, with supplementary values from the report’s retained data and chart generator. No live model calls.
- [1]Chance & name-association probes
report_evidence.pyandreport_data/. 350 calls; exact outcome means insupplementary_summary.json. - [2]Answer-order controls
report_controls.pyandreport_data/order_control_summary.json. 150 calls across five control conditions. - [3]Hiring & paired gender audit
jev_bias_analysis/README.md,FINDINGS.md, and itsdata/. Selected hiring chart values frommake_report_charts.py. - [4]Partisan trust experiment
quick_bias_probe/FINDINGS_partisan_trust_bias.md; raw data in itsdata/. 16 bills, 64 cells, 2,944 requests. - [5]Moral foundations & actor labels
quick_bias_probe/study_morality.py; raw data in itsdata/. Displayed scores inreport_data/morality_summary.json. - [6]Theory referencesGraham et al., “Mapping the Moral Domain,” JPSP 101 (2011), 366–385. doi:10.1037/a0021847 ↗
Iyer et al., “Understanding Libertarian Morality,” PLOS ONE 7 (2012), e42366. doi:10.1371/journal.pone.0042366 ↗
Repository paths above identify the original research artifacts; they are citations, not hosted download links. The new chance/name probes and controls used the versioned model ID shown at the top. Reported intervals and estimates retain the source analysis’s scope.