Preliminary baseline · PCT and WVS side by side. Seven of eight models switch the economic sign.
Amtro signal
Seven of eight models switch economic quadrants between two instruments.
On the Political Compass Test, all eight scorable models landed economically left and socially libertarian. On the Amtro WVS battery, seven of them crossed the economic zero line. A political label from a single test can shift the direction of the classification.
8scorable models
3survey instruments
7/8quadrant switches
2026-07baseline
Scroll horizontally on narrow viewports
Political Compass Test (PCT) and WVS-derived battery, July 2026 snapshot. Eight models scored on both instruments. Seven change economic sign; none change social sign. Scales are differently constructed — compare direction/sign, not absolute position. Status: provisional. Moonshot Kimi PCT: 1 of 3 planned runs. Planned runs otherwise: 3, temperature 0.
Protocol outputs, not beliefs or personality; no provider, country or training effect.
Data as table (text alternative)
PCT and WVS coordinates for eight models, July 2026
ID
model_id
PCT eco
PCT soc
WVS eco
WVS soc
eco switch
PCT n
A
anthropic/claude-sonnet-4.6
−8.4533
−8.4105
+1.0417
−2.4583
yes
3
B
deepseek/deepseek-v3.2
−8.0783
−7.4874
+1.0764
−2.1042
yes
3
C
z-ai/glm-5
−4.3700
−5.3849
+0.8472
−2.6667
yes
3
D
google/gemini-3.5-flash
−6.6200
−7.4874
+0.8264
−3.3194
yes
3
E
moonshotai/kimi-k2.5
−5.7450
−6.8208
+0.4514
−3.6250
yes
1
F
mistralai/mistral-large-2512
−8.7450
−7.8464
+1.4792
−2.5972
yes
3
G
openai/gpt-5.5
−6.5367
−7.7438
−0.1736
−4.3264
no
3
H
qwen/qwen3.6-plus
−7.3700
−7.8464
+0.2292
−3.8333
yes
3
Totals: economic sign changes 7/8; social 0/8. Status: provisional.
Observation
A compass score is not a stable model property.
Instrument, prompt, language, response format and model version all affect the measured position. Political model cards should disclose these conditions instead of treating a single point as a fixed orientation.
Next step
The baseline becomes a time series.
The series is set up for monthly repetition. Future runs will track model changes and drift; quarterly robustness checks will vary question order, language and open-ended response conditions.
Citation
Observation, inference and hypothesis stay separate.
Amtro Research (2026). Same model, different worldview? Instrument sensitivity of political LLM evaluations. Amtro AI Map, baseline July 2026, preliminary version.