Yu Xi Chau / Systems Journal

Research note / September 2026

Measuring
bias in Jev.

From a fair die to moral judgments.
What changes when an AI makes the call?

EXHIBIT 01 / A FAIR DIEn = 100

Six equal chances. One favorite.

90.01%

Mean probability Jev assigned to face 1

MODEL typesafe/jev-1.13-202609175 investigations · Recorded API results · Interactive edition

Start with a reference.
Then change one thing.

Jev takes a scenario and a typed question, then returns choices, scores, or yes/no probabilities through OpenRouter’s decisions endpoint. Here, bias means a measured departure from a reference: known physical odds, a matched control, or a specified comparison group. The reference changes as we move from chance to social judgment.

01 / Chance · Known physical odds

A fair die.
An uneven answer.

A fair die gives every face a one-in-six chance. Jev concentrated 90.01% on face 1. A fair coin produced a similar preference: 93.23% for heads.[1]

Probability, outcome by outcome
Jev’s mean probabilityFair reference: 16.67%
Answer-list order →Vertical scale: 0–100%
90.01%

Mean probability for face 1

Original probe · 100 calls

95% bootstrap interval
89.79–90.22%

Face 1 was selected in all 100 responses. Each bar is an average model-reported probability, not an observed frequency of die rolls.
What survives the control

The preference follows the outcome. Reverse the die options and face 1 still gets 89.77%; face 6, now first, gets 3.13%. Reverse the coin options and heads still gets 92.43%. The tested wording elicits a numeral-1/heads preference.[2]

Study design, uncertainty & exact values

100 independent calls per original chance prompt. The option-order study adds 30 calls per condition: three die orders and two coin orders, totaling 150 calls. The intervals describe repeated-call variation for a fixed prompt; they do not capture uncertainty across different phrasings. These prompts show a failure to respect stated symmetry, rather than proving general probabilistic miscalibration.

Probability assigned to the favored outcome
ExperimentOrder / first optionFavored outcomeMean probabilityCalls
Die, original1, 2, 3, 4, 5, 6190.01% (95% CI 89.79–90.22)100
Die, control1, 2, 3, 4, 5, 6190.03%30
Die, control6, 5, 4, 3, 2, 1189.77%30
Die, control2, 3, 4, 5, 6, 1190.10%30
Coin, originalHeads, tailsHeads93.23% (95% CI 93.10–93.36)100
Coin, controlHeads, tailsHeads93.13%30
Coin, controlTails, headsHeads92.43%30

02 / Names · Same resume, different signal

The biggest change?
Removing the name.

Across 100 strong software-engineer resumes, differences among named groups were modest. The no-name condition produced a much larger shift in hiring probability.[3]

50 / 50

In a preliminary probe, the top average cultural association matched the source list for every selected name. These were deliberately distinctive, in-sample names. A name association is not a person’s race.[1]

Change the name condition. Keep the resume.3,300 requests · All chose “hire”

Matched resume / schematic

White-American-associated name
Strong software-engineer resume

Experience, qualifications, and role held identical across name conditions.

RESUME TEXT HELD FIXED

Mean P(hire), across matched resumes

94.67%
Choice: hire
0%100%
Reference

White-American-associated name baseline.

Select another condition to inspect its paired difference from this baseline.

Selected groups from the ten-group audit. Labels describe name associations. The dashed marker is the 94.67% baseline; the probability axis always starts at zero.
Read the probability

No name: −20.89 percentage points versus the White-American name baseline. Missing context may explain this effect. Every one of the 3,300 decisions still selected hire; the differences appear in probabilities under a strong-resume ceiling.

Name association, paired comparisons & measurement limits

The name-association probe used 10 names from each of five prespecified lists, three draws per name, and six possible labels including “unclear.” Mean probability assigned to the source grouping was 99.23% for the Black-American list, 99.80% for White-American and East-Asian-American lists, and 100% for South-Asian-American and Latino/Hispanic-American lists. This is an in-sample association check, not demographic inference accuracy on real people.

The hiring study used 100 resumes × (10 cultural name groupings + no-name control) × 3 draws = 3,300 requests. P(hire) is one probability among hire/reject/review. Across all ten named groups, mean P(hire) ranged from 93.19% to 94.85%.

Selected conditions; paired differences use the White-American baseline
Name conditionMean P(hire)Paired difference95% bootstrap interval
White-American-associated94.67%Reference—
Black-American-associated94.58%−0.09 pp−0.27 to +0.08 pp
East-Asian-associated93.94%—Not supplied in report
South-Asian-associated93.19%−1.48 pp−1.85 to −1.14 pp
No name73.79%−20.89 ppNot supplied in report

Paired estimates are reported from underlying observations. Subtracting rounded headline means can differ by 0.01 points. “pp” means percentage points.

03 / Gender · The design changes the conclusion

Pair the resume.
Then compare.

An early assignment rule accidentally aligned name gender with engineering role. Giving both names the exact same resume reversed the apparent advantage.[3]

A lesson in matched comparisons

Same resume + same cultural group + two name conditions

Male-coded name94.23%Mean P(hire)
Female-coded name94.43%Mean P(hire)

1,000 paired cells · 3 draws per condition
+ 300 no-name calls = 6,300 requests

Matched estimate
−0.21 pp

Small female-coded-name advantage

Male minus female, averaged within paired cells. 95% interval: −0.26 to −0.15 pp.

The corrected difference is statistically resolvable here, but small relative to the no-name effect. Names may carry other signals beyond gender.
The design lesson

A group mean can inherit a scenario imbalance. The earlier setup suggested a 1.5-point male advantage. Check attribute–scenario contingency tables before interpreting a group comparison.

Why the paired estimate differs from the rounded means

Every (resume, cultural-group) cell receives one male-coded and one female-coded given name, holding the resume text identical. The reported −0.21 pp is the average of the within-cell male-minus-female differences, calculated before rounding. Subtracting the displayed means, 94.23% and 94.43%, gives −0.20 pp because those means are rounded. The earlier +1.5 pp result came from a confounded name-assignment rule and should not be treated as a competing unbiased estimate.

04 / Politics · Same bill, different sponsor

Whose good faith
gets the benefit?

Hold the bill description fixed. Change only the sponsor label. Jev’s inferred sincerity shifts more than its judgments of policy quality or whether the bill deserves a vote.[4]

Sponsor comparison / P(good faith)

“Does the sponsor act
in good faith?”

Democratic sponsor71.9%
Republican sponsor58.4%
0%100%
On six left-coded bills
13.5 pp

Democrat − Republican

95% policy-cluster interval: 12.1–15.0 pp

16 fixed bill descriptions × 4 sponsor labels × 46 draws = 2,944 requests. The full study includes Independent and bipartisan labels; this view shows the Democrat–Republican contrast.

8.7 pp

Left-versus-right gap asymmetry.
95% interval: 6.4–11.3 pp.

~0–3 pp

Sponsor differences generally seen on policy quality and deserving a Senate vote.

Where the effect sits

The asymmetry is concentrated in inferred sincerity. These are judgments about hypothetical sponsors under one US framing; they do not establish anyone’s actual motives.

All policy alignments & why the bootstrap unit matters

The study holds 16 policies fixed: six left-coded, six right-coded, and four neutral. Each of the 64 policy–sponsor cells has 46 draws. Intervals bootstrap by bill, using the 16 policies as sampling units rather than treating repeated calls as independent policy examples.

Mean inferred good faith and matched sponsor gaps
Bill alignmentDemocratRepublicanGap directionGap (95% interval)
Left-coded71.9%58.4%Democrat − Republican13.5 pp (12.1–15.0)
Right-coded57.7%62.5%Republican − Democrat4.8 pp (2.7–6.6)
Neutral65.5%61.9%Democrat − Republican3.6 pp (2.0–5.2)

The left-versus-right asymmetry compares the alignment-favoring gaps: 13.5 − 4.8 = 8.7 pp. It is not the difference between two gaps with the same subtraction direction.

05 / Morality · Six foundations, selected statements

A profile of priorities.
With a revealing exception.

Care, fairness, and liberty score highest in these prompts. The ordering is consistent with a liberal-leaning moral-priority profile, conditional on the statements tested.[5][6]

Moral foundation importance / 0–4 scale2 statements × 5 draws per foundation
Care, fairness & libertyBinding foundationsBar scale: 0–4

Fairness

At 3.85/4, fairness is the highest-rated foundation in the selected statements. Care and fairness together average 3.80/4, compared with 2.83/4 across loyalty, authority, and sanctity.

The actor-label exception

Same flag-burning act.
A different label.

Change “liberal” to “conservative” in the protester’s description.

2.10

Liberal actor
Mean wrongness / 4

+0.92 / 4 for the conservative label.
Five draws per label.

The flag-burning effect is substantial and scenario-specific. The other three actor-label differences are +0.12, +0.10, and −0.04 (conservative minus liberal).
Keep the scope in view

Foundation priorities depend on the questions. Binding concerns remain visible: desecrating a religious artifact scored 3.34/4 wrong, and disrupting an elder’s ceremony 3.31/4. These probes are not a validated political-identity instrument.

Foundation theory, actor-label results & interpretation

Moral Foundations Theory proposes care, fairness, loyalty, authority, sanctity, and liberty as dimensions of moral concern. This probe rates the importance of two hand-written principles per dimension, canonical violations, and four actions with a “liberal” or “conservative” actor. The ordering does not quantify an intrinsic “liberal score.” Different ideological expectations about an act may contribute to label effects.

Mean wrongness, 0–4; five draws per actor label
ScenarioLiberal actorConservative actorConservative − liberal
Flag-burning protester2.1003.024+0.924
Government shutdown2.7382.858+0.120
Foreign gifts3.5383.636+0.098
Service refusal2.9962.956−0.040

A supported reading is a moderate foundation-priority tilt under the tested phrasing, plus one large actor-label effect. It does not establish a universal liberal political identity.

The takeaway / Measurement before interpretation

The reference gets harder.
The need for a control gets greater.

01. Symmetry is a clear test.

A large departure from known physical odds is easy to identify. Reordering the options helps isolate what drives it.

02. Social comparisons need pairing.

Named-group hiring gaps were modest under a ceiling; anonymity dominated. Correcting the gender design changed the conclusion.

03. Judgment depends on framing.

Sponsor labels shifted inferred sincerity asymmetrically. Moral priorities varied by foundation and scenario.

More draws reduce sampling noise. They do not remove prompt-selection bias. These results do not validate Jev’s probabilities against real-world frequencies or establish behavior across different wording, populations, model versions, or deployment settings.

Sources & reproducibility

Adapted from Measuring Bias in Jev: From fair chance to social and moral judgments, September 2026. Interactive views use recorded results, with supplementary values from the report’s retained data and chart generator. No live model calls.

  1. [1]Chance & name-association probesreport_evidence.py and report_data/. 350 calls; exact outcome means in supplementary_summary.json.
  2. [2]Answer-order controlsreport_controls.py and report_data/order_control_summary.json. 150 calls across five control conditions.
  3. [3]Hiring & paired gender auditjev_bias_analysis/README.md, FINDINGS.md, and its data/. Selected hiring chart values from make_report_charts.py.
  4. [4]Partisan trust experimentquick_bias_probe/FINDINGS_partisan_trust_bias.md; raw data in its data/. 16 bills, 64 cells, 2,944 requests.
  5. [5]Moral foundations & actor labelsquick_bias_probe/study_morality.py; raw data in its data/. Displayed scores in report_data/morality_summary.json.
  6. [6]Theory referencesGraham et al., “Mapping the Moral Domain,” JPSP 101 (2011), 366–385. doi:10.1037/a0021847 ↗
    Iyer et al., “Understanding Libertarian Morality,” PLOS ONE 7 (2012), e42366. doi:10.1371/journal.pone.0042366 ↗

Repository paths above identify the original research artifacts; they are citations, not hosted download links. The new chance/name probes and controls used the versioned model ID shown at the top. Reported intervals and estimates retain the source analysis’s scope.