Amtro AI Map · Baseline July 2026

Same model,
different worldview?

Why political labels for language models depend on the questionnaire used.

Preliminary baseline · Methods series · Limitations stated explicitly

Eight scorable models, three instruments. Status: preliminary.

Dual panel: Political Compass and WVS battery for eight models
Preliminary baseline · PCT and WVS side by side. Seven of eight models switch the economic sign.
Amtro signal

Seven of eight models switch economic quadrants between two instruments.

On the Political Compass Test, all eight scorable models landed economically left and socially libertarian. On the Amtro WVS battery, seven of them crossed the economic zero line. A political label from a single test can shift the direction of the classification.

8scorable models
3survey instruments
7/8quadrant switches
2026-07baseline
Political Compass Test (PCT) and WVS-derived battery, July 2026 snapshot. Eight models scored on both instruments. Seven change economic sign; none change social sign. Scales are differently constructed — compare direction/sign, not absolute position. Status: provisional. Moonshot Kimi PCT: 1 of 3 planned runs. Planned runs otherwise: 3, temperature 0.

Protocol outputs, not beliefs or personality; no provider, country or training effect.

Data as table (text alternative)
PCT and WVS coordinates for eight models, July 2026
ID model_id PCT eco PCT soc WVS eco WVS soc eco switch PCT n
Aanthropic/claude-sonnet-4.6−8.4533−8.4105+1.0417−2.4583yes3
Bdeepseek/deepseek-v3.2−8.0783−7.4874+1.0764−2.1042yes3
Cz-ai/glm-5−4.3700−5.3849+0.8472−2.6667yes3
Dgoogle/gemini-3.5-flash−6.6200−7.4874+0.8264−3.3194yes3
Emoonshotai/kimi-k2.5−5.7450−6.8208+0.4514−3.6250yes1
Fmistralai/mistral-large-2512−8.7450−7.8464+1.4792−2.5972yes3
Gopenai/gpt-5.5−6.5367−7.7438−0.1736−4.3264no3
Hqwen/qwen3.6-plus−7.3700−7.8464+0.2292−3.8333yes3

Totals: economic sign changes 7/8; social 0/8. Status: provisional.

Observation

A compass score is not a stable model property.

Instrument, prompt, language, response format and model version all affect the measured position. Political model cards should disclose these conditions instead of treating a single point as a fixed orientation.

Next step

The baseline becomes a time series.

The series is set up for monthly repetition. Future runs will track model changes and drift; quarterly robustness checks will vary question order, language and open-ended response conditions.

Citation

Observation, inference and hypothesis stay separate.

Amtro Research (2026). Same model, different worldview? Instrument sensitivity of political LLM evaluations. Amtro AI Map, baseline July 2026, preliminary version.

Download audit package (ZIP, 20 KB) Curated canonical baseline files (claim matrix, quality report, aggregate, protocol, lab and dual figures) with SHA-256 — not a full raw-data archive. Browse individual files

Updates

Hear about the next measurement.

The series is set up for monthly repetition. We only write when new results are published. No marketing; reply once to unsubscribe.