Shav Vimalendiran
BlogWiki ↗
  • ›Cultural Alignment of Open-Weight LLMs on the Inglehart-Welzel MapAug 4, 2026
  • The Last MoatsJul 7, 2026
  • The Tech FrontierJul 6, 2026
  • The Trust BarrierJul 4, 2026
  • I Open-Sourced My Coding Agents' MemoryJun 24, 2026
  • How I Gave My Coding Agents Persistent MemoryMar 12, 2026
  • Multi‑Agent Web Exploration with Shared Graph MemoryFeb 19, 2026
  • Our Lessons from Building Production Voice AIJan 11, 2026
  • Reinforcement Learning, Memory and LawDec 10, 2025
  • Automating Secret ManagementOct 23, 2025
  • The Knowledge LayerOct 5, 2025
  • OTel Sidecars on FargateSep 11, 2025
  • Git Disasters and Process DebtSep 7, 2025
  • Is Code Rotting Due To AI?Sep 3, 2025
  • The Integration IllusionAug 30, 2025
  • When MCP FailsAug 26, 2025
  • Context EngineeringAug 22, 2025
  • Stop Email Spoofing with DMARCAug 5, 2025
  • SOTA Embedding Retrieval: Gemini + pgvector for Production ChatJul 21, 2025
  • Agentic Design PatternsJun 21, 2025
  • Building AI Agents for Automated PodcastsJan 1, 2025
  • Rediscovering CursorDec 2, 2024
  • GraphRAG > Traditional Vector RAGAug 8, 2024
  • Cultural Bias in LLMsJul 20, 2024
  • Mapping out the AI Landscape with Topic ModellingJul 7, 2024
  • Sustainable Cloud Computing: Carbon-Aware AIJun 27, 2024
  • Defensive Technology for the Next Decade of AIJun 24, 2024
  • Situational Awareness: The Decade AheadJun 13, 2024
  • Mechanistic Interpretability: A SurveyJun 7, 2024
  • Why I Left UbuntuMay 24, 2024
  • Multi-Agent CollaborationApr 16, 2024
  • Building Better Retrieval SystemsMar 28, 2024
  • Building an Automated Newsletter-to-Summary Pipeline with Zapier AI Actions vs AWS SES & LambdaFeb 3, 2024
  • Local AI Image GenerationDec 15, 2023
  • Deploying a Distributed Ray Python Server with Kubernetes, EKS & KubeRayNov 15, 2023
  • Making the Switch to Linux for DevelopmentOct 24, 2023
  • Scaling Options Pricing with RayOct 1, 2023
  • The Async Worker PoolSep 23, 2023
  • Browser Fingerprinting: Introducing My First NPM PackageSep 8, 2023
  • Reading Data from @socket.io/redis-emitter without Using a Socket.io ClientJul 6, 2023
  • Socket.io Middleware for Redux Store IntegrationJul 1, 2023
  • Sharing TypeScript Code Between Microservices: A Guide Using Git SubmodulesApr 21, 2023
  • Efficient Dataset Storage: Beyond CSVsFeb 3, 2023
  • Why I switched from Plain React to Next.js 13Nov 8, 2022
  • Deploy & Scale Socket.io Containers in ECS with ElasticacheNov 3, 2022
  • Implementing TOTP Authentication in Python using PyOTPSep 13, 2022
  • Simplifying Lambda Layer ARNs and Creating Custom Layers in AWSSep 9, 2022
  • TimeScaleDB Deployment: Docker Containers and EC2 SetupJun 23, 2022
  • How to SSH into an EC2 Instance Using PuTTYDec 16, 2021
Loading post…

In This Post

1. Introduction2. Related Work3. Method3.1 The instrument and the survey data3.2 The projection pipeline, and why each step is there3.3 Administering the survey to models3.4 Uncertainty3.5 Refusals, and the selection they induce4. Results4.1 The corrected 2024 cohort4.2 The 2026 frontier generation, both arms5. DiscussionLimitationsEthics StatementAppendix A - ModelsAppendix B - The ten IVS items and the elicitation protocolAppendix C - Rotation sensitivity and reproducibilityFootnotes
Published: August 4, 2026
Next

Cultural Alignment of Open-Weight LLMs on the Inglehart-Welzel Map

PDFGithub

Large language models are increasingly embedded in workflows where cultural values shape outcomes, yet the values they express are rarely measured against established cross-cultural instruments - and where they have been, the measurement pipelines themselves have gone unaudited. We map open-weight LLMs onto the Inglehart-Welzel Cultural Map by administering the ten Integrated Values Surveys (IVS) items that define the map and projecting model responses through the same frozen pipeline as the 109 surveyed countries. Auditing our own earlier public analysis (July 2024) surfaced three defects that this paper corrects and quantifies: model responses were projected without the fitted standardisation and under a re-fitted rotation, so model and country points did not provably share a coordinate space; an SPSS user-missing code was read as data in 31.9% of survey rows, silencing one of the ten instrument items; and responses elicited in English and Chinese were pooled across a gap of more than ten sampling standard deviations. With the corrected pipeline we report eleven 2024 model-language cells with bootstrap confidence regions and two independent region-assignment rules, and find the substantive conclusion strengthened - among models that answer the instrument: every one of the eleven cells, of every origin, sits in the secular/self-expression quadrant in every bootstrap replicate; its mean position lies farther from the pooled human respondent mean than at least 95% of the 109 surveyed countries and never within 1.6 map units of a non-Western region centroid; and even the most adverse replicate of any cell stays farther than 82% of countries and 1.2 units clear. Elicitation language, not developer origin, is the largest factor moving a model on the map: the one 2024 model administered in both languages moves 1.77 map units between its English and Chinese runs. A 2026 factorial re-run over the cloud-served frontier generation (17 models × two administration languages) then tests these findings where the 2024 study could not look. Frontier Chinese-origin models now answer the instrument reliably (10/10 coherent, vs 4/9 in 2024), and - still among models that answer, though the one excluded cell's worst-case refusal bound also stays inside - every usable cell again lies in the secular/self-expression quadrant in all of its bootstrap replicates. But under English administration the homogenisation attenuates: the English arm sits markedly closer to the human mean than the 2024 English-administered cells (median distance 2.22 vs 2.71 map units, with 14 of 17 cells closer than the closest 2024 English cell) while the Chinese arm does not, and the English arm's spread of distances to the human mean collapses to a quarter of 2024's. With language a designed factor, the language effect replicates in kind - within-model displacements of 0.18–1.51 map units (mean 0.65 [0.52, 0.81]), multiples of the ~0.1 sampling SDs - and the traditional-pole shift vanishes: where the 2024 bilingual model moved sharply toward the traditional pole under Chinese administration, only 5 of 16 frontier models move that way at all, 15 of 16 move toward self-expression, both origin cohorts move away from the Confucian centroid, and Chinese- and Western-origin models are affected indistinguishably at this design's resolution (permutation p = 0.89) - the largest single displacement belongs to a French lab's model. Item-level structure suggests political-context gating rather than value shift: of the items that move under Chinese administration, the only shift toward the survival pole is petition-signing's. The Confucian hypothesis fails its first genuine test, in both languages. We release the corrected, validated, unit-tested pipeline and the complete 2026 response corpus, including refusals and reasoning traces.

Keywords: cultural bias, cultural alignment, values surveys, Inglehart-Welzel, evaluation, measurement validity

1. Introduction

Culture shapes causal attribution, judgement, trust in automation, privacy expectations, and health behaviour. As LLM-generated text becomes a transmission channel for language itself, models that systematically overrepresent one culture's values can propagate those values into every downstream context they touch. Whether that risk is real depends on an empirical question: what values do deployed models actually express, measured on instruments designed for measuring values?

An earlier version of this analysis was released publicly in July 2024 1 and has since been cited repeatedly in peer-reviewed venues - including the Columbia study whose language-resource result now supplies a mechanism for our central finding 2. This paper is what became of it when it was held to peer-review standards: a full audit of the measurement pipeline, three corrections, uncertainty on every claim, and a designed experiment where the original had a confound. We state the defects openly - they are quantified in §3 and each one changes something a reader of the original would have believed - because the paper's central methodological argument is that instrument-based claims about model values are only as good as the measurement code behind them, and that argument applies to our own work first.

The three audit findings, in brief.

  1. Projection defects. The 2024 code projected raw 1–10 model answers onto principal axes fitted to standardised survey data, and re-fitted the varimax rotation on the model scores. An uncentred projection is dominated by whichever items have the largest raw means - the ten IVS items have wildly different native ranges (trust is 1–2; importance of God is 1–10) - which pulls every model toward one location largely regardless of its answers; a second rotator fit places models in a different, incomparable coordinate space from the countries. Under the correction, one model (llama3:70b) initially moved by 2.05 map units - its published near-Catholic-Europe position was an artefact.
  2. A silenced instrument item. The IVS ships SPSS user-missing codes; -3 on the autonomy index Y003 was read as a datum in 31.9% of post-2005 rows (structurally: 15 countries carry it in 100% of rows, 65 in none). The item - one of the five that officially define the traditional/secular axis - consequently loaded ≈0 on both axes, and the "at least 6 of 10 items answered" filter counted a non-answer as an answer. Recoding and refitting raises explained variance from 38.4% to 41.7% and restores Y003 to a substantive loading (0.19, 0.24) with its expected correlational profile (r = −0.39 with importance of God).
  3. Pooled elicitation languages. The two models assigned to the Confucian region in 2024 were administered only in Chinese; the one model administered in both languages had its two runs pooled under one label, averaging across a 1.77-map-unit gap whose pooled row carries an implied error bar of ±0.1, both figures computed on the corrected pipeline (the 2024 post itself reported no uncertainty at all). Language is an experimental factor and is treated as one throughout.

Contributions.

  1. Methodological. A corrected, validated, unit-tested, open pipeline: full IVS reconstruction (EVS 1981–2017 + WVS 1981–2022) with probabilistic PCA retaining the 51% of rows (199,926 of 391,630) listwise deletion would discard; a single stored rotation applied identically to countries and models; survey-design weights applied at aggregation; two bootstrap estimators (item-level, which omits cross-item covariance, and clustered over prompt variants where the design permits); two region-assignment rules of different type reported with the classifier's own cross-validated accuracy; and classifier-free headline statistics with replicate-simultaneity handled explicitly. To our knowledge no prior instrument-based evaluation of LLMs reports confidence regions for model positions (§2). The differentiation from the closest prior work (Tao et al. 3) is drawn in §2: published country scores versus a full individual-level reconstruction, and a handful of API models versus Chinese-developed, Chinese fine-tuned, and alignment-stripped families under a language-controlled protocol.
  2. Empirical. An uncertainty-aware map of the 2024 open-weight generation as eleven model-language cells: all far from the human mean in the secular/self-expression direction, with elicitation language the largest single mover - a finding the 2024 analysis mistook for a Confucian signal.
  3. Experimental. A factorial language × model design over the 2026 cloud-served frontier generation (17 models, 10 Chinese-origin; 34 cells, 17,000 calls), turning the 2024 language confound into a designed within-model contrast - and giving the Confucian hypothesis its first genuine test, under both English and native-language administration. It fails in both arms: what little Confucian proximity exists appears under English administration and is not exclusive to Chinese-origin models, native-language administration moves both cohorts away from the Confucian centroid, and the origin × language interaction is null. In 2024 this test was impossible - five of nine attempted Chinese-origin models produced no parseable responses; in 2026 all ten parse at ≥96% in both languages, and the disappearance of that limitation is itself a finding (§4.2).

2. Related Work

Values-survey instruments. Tao, Viberg, Baker and Kizilcec 3 - now published in PNAS Nexus - is the founding work in this line: administering WVS items to GPT-family models, they find cultural values resembling English-speaking and Protestant European countries and show cultural prompting can shift them. Eren et al. 4 extend the cultural-prompting result to open-weight models, so open-weight coverage alone is no longer novel; our differentiation is measurement auditability - full individual-level IVS reconstruction rather than published country scores, one audited projection path shared by countries and models, uncertainty on every claim, and coverage of Chinese-developed, Chinese fine-tuned, and alignment-stripped families under a language-controlled protocol. LLM-GLOBE 5 evaluates Chinese- and US-developed models on the GLOBE dimensions and finds significant China-US model differences on seven of nine dimensions - but also that both cohorts diverge significantly from their own countries' human ground truth, which is the homogenisation pattern we measure on the IW instrument. Luther and Brown 6 find DeepSeek expressing US-aligned values on Hofstede's VSM13 under every prompting strategy tried - converging evidence, by an independent instrument, that Chinese-developed frontier models do not express Chinese-aligned values.

Elicitation-language effects. Bulté and Rigouts Terryn 7 probe ten LLMs with Hofstede/WVS items in eleven languages and independently establish that prompt language shifts expressed values while models remain anchored to a small set of cultural defaults - though in their factor ordering explicit cultural framing outweighs prompt language, and the magnitude of language effects is genuinely contested (a direct replication attempt of a flagship prompt-language result reports near-zero effects 8). Two clarifications position our claim precisely: ours is a different comparison - elicitation language against developer origin, not against cultural framing - and it is a per-model quantity: our factorial design estimates δm\delta_mδm​ - a model's positional displacement between its Chinese- and English-administered runs - for each model individually (realised for 16 of the 17 - the voided nemotron-3-ultra Chinese cell leaves no pair) rather than asserting a universal effect, because the literature indicates the effect is strongly model-dependent. Our contribution is therefore not the existence of the effect but its magnitude relative to sampling uncertainty on a validated instrument - 1.77 map units against per-cell SDs of ~0.1 - and the demonstration that our own 2024 analysis's Confucian-region assignments rest entirely on cells where elicitation language and developer origin are perfectly confounded. Kazemi et al. 2 supply the most direct mechanistic account: the volume of online language resources predicts how faithfully a model represents a country's WVS values (explaining 44% of variance in GPT-4o, with error rates five times higher in the lowest-resource languages) - under that account, both the homogenisation we observe and the language-of-administration shift are footprints of resource-weighted training distributions.

Survey validity for LLM respondents. Rupprecht, Ahnert and Strohmaier 9 run 167k simulated WVS interviews and document human-like response biases - including the central-tendency behaviour we quantify on the IW projection (§5) - plus item nonresponse from censored models, anticipating our refusal analysis. Himelstein et al. 10 show refusal masks rather than removes value-laden bias, which upgrades our "conditional on answering" limitation from a selection caveat to a mechanism claim. Li, Li and Qiu 11 find silicon samples homogenised toward majority views; Taday Morocho et al. 12 find persona conditioning degrades alignment with real respondents, supporting our choice not to persona-condition beyond the neutral "average human" framing.

The broader elicitation toolkit. Scenario-based probing and steering 13 with latent activation steering 14 (one research programme), culturally grounded personas 15, word-association tests 16, and multi-agent cultural debate 17 form a growing toolkit for eliciting and shifting cultural behaviour, several of them working directly on the IW axes. Against that toolkit our contribution is deliberately narrow and instrument-first: a decades-validated survey instrument, administered under a controlled protocol, with the measurement pipeline itself audited, validated and released.

On linguistic relativity: the Sapir-Whorf tradition 18 motivates the possibility that language alone shapes elicited values. Our language-of-administration finding is consistent with a weak form of that claim, though the design cannot separate what the prompt language activates in the model from what the training data associated with that language contains - and per Au's critique of the strong form, we treat the mechanism as an open question, not a conclusion.

Where this paper sits. The field now has three kinds of paper: elicitation papers, which propose new ways to surface cultural behaviour (personas, scenarios, debate, word association); steering papers, which move it (prompting, activation steering); and measurement papers, which ask what models express on validated instruments. Ours is a measurement paper with a specific thesis the other two kinds presuppose but rarely check: that the measurement itself is right. Its distinguishing artefacts are the audit (three quantified defects in our own prior pipeline, one of which - elicitation-language pooling - also implicates published claims built on unaudited elicitation), the uncertainty machinery (to our knowledge, no prior instrument-based evaluation reports confidence regions for model positions), and the released, tested, data-free-CI pipeline that lets any reviewer re-derive every number. Where elicitation and steering papers show what models can express, we pin down what they do express by default, with error bars - the quantity a deployment inherits when nobody intervenes.

3. Method

3.1 The instrument and the survey data

The Inglehart-Welzel (IW) map 19 places societies on two axes - survival vs. self-expression and traditional vs. secular-rational values - derived from ten IVS items: feeling of happiness (A008, 1–4), interpersonal trust (A165, 1–2), respect for authority (E018, 1–3), petition signing (E025, 1–3), importance of God (F063, 1–10), justifiability of homosexuality (F118, 1–10) and abortion (F120, 1–10), national pride (G006, 1–4), the post-materialist index (Y002), and the child-qualities autonomy index (Y003, −2…2). We merge the EVS Trend File 1981–2017 20 and the WVS Trend 1981–2022 21 with the published merge syntax, filter to waves from 2005 onward, recode out-of-range SPSS user-missing sentinels to missing before any completeness filtering (the recode is range-driven per item, so any sentinel - not only the Y003 -3 that bit the 2024 analysis - is caught), and require at least six of the ten items per respondent: 391,630 individual responses across 109 countries and territories. Country-level aggregation applies the IVS equilibrated weight (S017), which corrects within-country sampling design; unweighted means would treat every respondent as equally representative of their country.

Missingness is largely by design: items rotate in and out of country-waves, so which entries are missing is determined by the survey administration, not by the respondent's values. This is the benign case for model-based imputation - the missing-at-random assumption is unusually plausible when missingness is planned 22 - and it is why we can retain the 51% of rows (199,926 of 391,630) a complete-case analysis would discard. Figure 1 shows the country map the fitted projection (§3.2) reconstructs.

The Reconstructed Instrument

All 109 countries and territories, projected from the IVS microdata by the corrected pipeline. No WVS figure is reproduced - the familiar regional clustering emerges from our own fit.

⌘ + scroll to zoom · drag to pan
× marks the pooled human respondent mean at (0.38, −0.01) - the reference point for every distance statistic in the paper.

Figure 1. The reconstructed instrument. No WVS-copyrighted figure is reproduced anywhere in this paper - this map is computed from the IVS microdata by our own pipeline, which doubles as a face-validity check: the familiar regional clustering emerges from our fit.

3.2 The projection pipeline, and why each step is there

Why not listwise deletion, and why not marginal imputation. Mean or marginal imputation operates at column level and destroys the cross-item covariance structure the components are estimated from; listwise deletion selects on the survey design (whole country-waves would vanish). Probabilistic PCA 23 is the estimator that does neither: it treats the data as generated by a two-dimensional Gaussian latent variable, fits by maximum likelihood - maximising the observed-data log-likelihood L(θ)=∑ilog⁡f(xiobs∣θ)\mathcal{L}(\theta) = \sum_i \log f(\mathbf{x}_i^{\mathrm{obs}} \mid \theta)L(θ)=∑i​logf(xiobs​∣θ) over the observed entries only, which has no closed form and is maximised by EM - and imputes missing entries from the joint model rather than from the margin. Our implementation is derived from pca-magic 24, seeded, capped (EM raises rather than silently returning an unconverged fit), and unit-tested against synthetic low-rank ground truth. The two retained components explain 41.7% of the variance (denominator: the total variance of the EM-completed standardised matrix, which conditional-mean imputation shrinks slightly below the nominal ten; the pre-recode 38.4% uses its own EM-completed denominator, so the two figures are each internally consistent but not a fixed-denominator comparison).

Standardisation. Each item is z-scored with parameters frozen at fit time:

zj=xj−μjσj,j=1,…,10z_j = \frac{x_j - \mu_j}{\sigma_j}, \qquad j = 1,\dots,10zj​=σj​xj​−μj​​,j=1,…,10

Standardisation is not cosmetic here. PCA finds directions of maximal variance, so if one feature varies more than the others only because of its scale, it dominates the components - and the ten IVS items natively span ranges from 1–2 (trust) to 1–10 (importance of God, justifiability items). Projecting unstandardised model answers onto axes fitted to standardised survey data - the 2024 defect - therefore let the wide-range items dominate every model's position regardless of its actual answer profile. That is the mechanism behind the pattern the audit predicted and found: under the correction most models barely moved, while the model whose position had been most propped up by raw-scale dominance (llama3:70b) moved 2.05 map units.

Components. With standardised data, the fitted covariance admits an orthonormal eigenbasis (it is symmetric, so all eigenvalues are real and eigenvectors orthogonal):

Σ=CΛC⊤,C⊤C=I,Λ=diag⁡(λ1,λ2),λ1≥λ2\Sigma = C \Lambda C^{\top}, \qquad C^{\top} C = I, \qquad \Lambda = \operatorname{diag}(\lambda_1, \lambda_2), \quad \lambda_1 \ge \lambda_2Σ=CΛC⊤,C⊤C=I,Λ=diag(λ1​,λ2​),λ1​≥λ2​

The projection matrix is just the selected eigenvectors concatenated - the columns of CCC - so scores are S=ZCS = ZCS=ZC with cov⁡(S)=Λ\operatorname{cov}(S) = \Lambdacov(S)=Λ diagonal: the two raw components are uncorrelated before rotation.

Rotation - fitted once, stored, disclosed. A varimax rotation is fitted once, on the training score matrix, and stored as an explicit orthogonal matrix

R(θ)=(cos⁡θ−sin⁡θsin⁡θcos⁡θ),R⊤R=I,θ^=−39.6°R(\theta) = \begin{pmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{pmatrix}, \qquad R^{\top} R = I, \qquad \hat{\theta} = -39.6°R(θ)=(cosθsinθ​−sinθcosθ​),R⊤R=I,θ^=−39.6°

Three disclosures a reviewer should want. (i) The criterion. Fitting varimax to the N×2N \times 2N×2 score matrix with Kaiser row-normalisation is not the textbook loadings rotation; it is in the family of the score-side varimax of Rohe & Zeng 25, with two acknowledged deviations from their estimator: we rotate unwhitened rather than whitened scores (the whitened variant agrees to 0.5°, Appendix C), and their consistency guarantees assume leptokurtic factor scores, which a ten-item ordinal cloud satisfies only approximately. We adopt it because its output matches the published IW orientation - textbook varimax on the 10×210 \times 210×2 loadings gives −23° to −27° and visibly rotates the map away from the published layout - an output-referenced choice we make transparent by releasing the full sensitivity grid (six criterion variants spanning −3.8° to −39.6°, Appendix C) rather than presenting it as inevitable. (ii) Kaiser normalisation is load-bearing: the same scores without row-normalisation rotate by only −3.8°. (iii) The rotated axes are correlated. By bilinearity of covariance, cov⁡(SR)=R⊤ΛR\operatorname{cov}(SR) = R^{\top} \Lambda Rcov(SR)=R⊤ΛR, which is diagonal only if Λ∝I\Lambda \propto IΛ∝I; with our eigenvalues and θ^\hat\thetaθ^ this gives r = 0.37 between the plotted axes. Orthogonal rotation of unwhitened principal scores does not preserve uncorrelatedness, no claim in this paper depends on the axes being uncorrelated, and readers eyeballing "independent" axes on the map should know they are not.

The rotation's remaining sign/order ambiguity is pinned by two anchor items with unambiguous placement in the IW literature: F118 (justifiability of homosexuality) loads positive on self-expression, F063 (importance of God) negative on secular-rational.

Rescaling. The published WVS constants (PC1′=1.81⋅PC1+0.38PC1' = 1.81 \cdot PC1 + 0.38PC1′=1.81⋅PC1+0.38, PC2′=1.61⋅PC2−0.01PC2' = 1.61 \cdot PC2 - 0.01PC2′=1.61⋅PC2−0.01) presuppose unit-variance factor scores, so rotated scores are standardised by their fitted standard deviations sss before the constants are applied - omitting this (as the 2024 analysis did) inflates the absolute scale by the fitted score SDs themselves, sss = 1.45/1.35 per axis (+45%/+35%), which matters because region centroids and all distance statistics live on the absolute scale. The full frozen path for any projected point - survey respondent, country aggregate, or model - is therefore

position=(x−μσ)CRs⋅a+b\text{position} = \frac{\left(\frac{\mathbf{x} - \boldsymbol{\mu}}{\boldsymbol{\sigma}}\right) C R}{s} \cdot a + bposition=s(σx−μ​)CR​⋅a+b

with every parameter fixed at fit time, read elementwise where shapes differ: x,μ,σ\mathbf{x}, \boldsymbol{\mu}, \boldsymbol{\sigma}x,μ,σ are 10-vectors (the projected point's raw item values and the frozen item means and SDs), CCC is the 10×210 \times 210×2 loading matrix, RRR the stored rotation, and s=(1.45,1.35)s = (1.45, 1.35)s=(1.45,1.35), a=(1.81,1.61)a = (1.81, 1.61)a=(1.81,1.61), b=(0.38,−0.01)b = (0.38, -0.01)b=(0.38,−0.01) are per-axis constants. (Distances throughout are Euclidean on this plane; these distances are the "map units" of every headline number.) Nothing downstream of the survey fit is ever re-estimated when new data is projected. That single sentence is the entire correction of 2024's defect #1, and the pipeline enforces it structurally: there is one project() method, it takes raw responses, and countries and models both go through it.

Validation (what each check proves, and what it does not). Three released checks gate the pipeline:

  1. Path identity (exact). All complete-case survey rows are pushed through the public projection API - the same function models go through - and compared with the coordinates the fit assigned them internally: maximum absolute discrepancy 0.0, at machine precision, over 191,704 rows. This is a regression test that the model path and the country path are byte-identical - precisely the property the 2024 code lacked, in both its unstandardised projection and its re-fitted rotation. It is not, and we do not claim it is, evidence that the map itself is right: both paths could share a wrong rotation and this check would still pass. Its value is that the class of defect that produced the 2024 error can never silently return.
  2. Correction accounting. Per-axis affine regression of the 2024 published coordinates on the corrected ones: R² of 0.990 and 0.993 with slopes 1.36/1.31 - the 2024 map was 36% and 31% wider per axis respectively, which is the intended effect of the unit-variance rescaling (the fitted slopes sit slightly below the theoretical 1.45/1.35 because the Y003 recode perturbs the regressor - an errors-in-variables attenuation; r > 0.99 per axis). This is an internal consistency check on the size of our own correction, not external validation: no regression against the published WVS country coordinates is performed, and the corrected map is the same map, moved by exactly the corrections we made and nothing else.
  3. Interpretability. The rotated loadings recover the published item structure - the justifiability items and the autonomy index load positive on the self-expression/secular directions, while God, pride and authority mark the traditional pole (F063 loads negative on the secular axis; G006 and E018 are reverse-keyed scales - 1 = proudest, 1 = greatest respect - so their raw loadings are positive) - including Y003, which the sentinel bug had silenced.

Beyond the three gates, a 20-seed refit harness bounds Monte Carlo noise from the EM initialisation (rotation span 5.2°; country coordinates mean SD 0.02, worst-case range 0.26 map units - same order as some displacements we report, hence included in the uncertainty budget; full budget in Limitations and data/seed_sensitivity.csv), and the merge and EM stages assert their invariants (row-count preservation, convergence) rather than degrading silently.

3.3 Administering the survey to models

Protocol. Each model receives each item under ten system-prompt persona variants ("You are an average human being responding to the following survey question…", nine near-paraphrases), five repeats each - 500 calls per model-language cell. Strict per-item parsers validate responses (bare integer in range; an ordered pair for Y002; up to five unique choices for Y003), with up to 15 (2024) / 3 (2026) re-asks on parse failure and a trailing "Sure thing! Here is my numerical answer:" primer as refusal mitigation. Both the retry policy and the primer are themselves measurement choices with consequences we return to in §3.5 and Limitations.

Language as a factor. The 2024 corpus decomposes into eleven model-language cells: seven models English-only, two Chinese-only, one in both languages (analysed as two cells). The 2026 design crosses all 17 models with both languages - 34 cells of 500 calls each, collected to design completeness (every item × variant × repeat key present exactly once; zero duplicate keys) - with the Chinese arm using corrected translations (the 2024 Chinese F118 prompt labelled both scale poles "always justifiable", and several system-prompt variants collapsed to duplicates in translation; both defects are documented in Appendix B). The within-model language contrast δm=positionzh−positionen\delta_m = \text{position}_{zh} - \text{position}_{en}δm​=positionzh​−positionen​ is therefore estimated under an otherwise identical protocol, for every model, in both directions of the design.

What gets recorded. The 2024 harness stored only (model, item, parsed_value) - which is why prompt-level variance could not be decomposed retrospectively and why refusals left no trace. The 2026 harness records raw response text, the separated reasoning trace for thinking models, the system-prompt variant id, the repeat index, per-attempt errors, and latency. Refusals are data.

3.4 Uncertainty

Why the point estimate forgives the 2024 pseudo-respondent construction - and the variance does not. The 2024 analysis assembled stored responses into ~50 "pseudo-respondents" per model by arbitrary pairing before projection. Because the frozen pipeline is affine in the item values, a model's mean position is an affine function of its per-item mean responses,

p=w0+∑j=110wjxˉj,p = w_0 + \sum_{j=1}^{10} w_j \bar{x}_j,p=w0​+j=1∑10​wj​xˉj​,

so the pairing never affected where a model landed. But covariance is a bilinear operator - Cov⁡[aX+b,cY+d]=acCov⁡[X,Y]\operatorname{Cov}[aX+b, cY+d] = ac\operatorname{Cov}[X,Y]Cov[aX+b,cY+d]=acCov[X,Y], Cov⁡[X+Y,Z]=Cov⁡[X,Z]+Cov⁡[Y,Z]\operatorname{Cov}[X+Y,Z] = \operatorname{Cov}[X,Z] + \operatorname{Cov}[Y,Z]Cov[X+Y,Z]=Cov[X,Z]+Cov[Y,Z] - and iterating those properties over the affine map gives

Var⁡(p)=∑j=110wj2Var⁡(xˉj)+2 ⁣ ⁣∑j<kwjwkCov⁡ ⁣(xˉj,xˉk).\operatorname{Var}(p) = \sum_{j=1}^{10} w_j^2 \operatorname{Var}(\bar{x}_j) + 2\!\!\sum_{j<k} w_j w_k \operatorname{Cov}\!\left(\bar{x}_j, \bar{x}_k\right).Var(p)=j=1∑10​wj2​Var(xˉj​)+2j<k∑​wj​wk​Cov(xˉj​,xˉk​).

The affineness argument says nothing about the second sum. All ten items were elicited under the same ten system-prompt variants, so a variant that pushes a model toward, say, more agreeable answers moves several items together: the cross-item covariances are not zero, and a bootstrap that resamples items independently forces them to zero (zero covariance does not imply independence; and resampling under an independence assumption does not make the true cross-item covariances zero - it only removes them from the variance estimate).

Two bootstraps, reported as what they are. The item bootstrap (B = 1,000) resamples each item's stored responses independently with replacement and projects the per-item means; by the identity above it sets every cross-item covariance to zero. With mixed-sign item weights the neglected covariance term can take either sign, so this is not a guaranteed bound: empirically it understates the cluster bootstrap on 27 of 33 2026 cells and overstates it on at least one axis for the other six. The cluster bootstrap (B = 10,000) resamples the ten prompt variants with replacement, carrying all items and repeats within a variant, and so propagates prompt-level correlation into the position; it is the primary estimator wherever the variant identifier was recorded. On the completed 2026 corpus, per-item intraclass correlations by prompt variant reach 1.00 (degenerately, under deterministic decoding; median 0.10; 51 of 340 cell-item pairs, over all 34 collected cells × 10 items, exceed 0.5 - for gemma4:31b's Chinese-administered F063, every repeat within a variant is identical while variants disagree, the prompt-variance-dominant regime in its purest form), and cluster-bootstrap regions are substantially larger than item-bootstrap ones - the product of a cell's two axis SDs is a median 2.9× greater (range 0.97–5.8× across the 33 usable cells) - routinely the difference between "significantly apart" and "not". With five repeats per variant, the top intraclass correlations imply design effects up to 5.0 - in those cells a cell-item mean's fifty calls carry the information of roughly ten independent draws, one per prompt variant (at the median ICC of 0.10 the design effect is 1.4). Determinism drives the gap: the cells with the smallest item-bootstrap SDs are precisely the most deterministic ones (Spearman ρ = 0.65 between mean per-item response entropy and item-bootstrap SD, p < 10⁻⁴ treating the 33 cells as independent; at model level, n = 17, the association holds at p ≈ 0.005; a standalone test outside every correction family), while the cluster SD is essentially unrelated to determinism - under near-deterministic decoding, token-sampling variance stops measuring anything a reader cares about, and the prompt-variant bootstrap must carry the uncertainty budget. The 2024 corpus did not record the variant id, so all 2024 ellipses are item-bootstrap ellipses omitting the cross-item term, with the 2026 ratios as the empirical scale of the likely understatement. With only ten clusters the cluster bootstrap is itself approximate 26.

Language displacement. The within-model contrast (§3.3) is summarised as δ^m\hat{\delta}_mδ^m​, the difference of the two arm mean positions, and its magnitude as the plug-in norm ∥δ^m∥\|\hat{\delta}_m\|∥δ^m​∥. The pairing of replicate iii in one arm with replicate iii in the other is arbitrary - the arms' bootstrap draws are independent - and the plug-in point estimate is invariant to it; the mean of per-replicate norms, by contrast, is a folded statistic, upward-biased wherever the true displacement is small against sampling noise, and is carried in the released artefacts as the alternative. Confidence intervals for ∥δm∥\|\delta_m\|∥δm​∥ are percentile intervals of the per-replicate norms; being intervals on a folded statistic they sit slightly above the plug-in point estimate wherever displacement is small against noise, and because a norm is non-negative they cannot be used as a test against zero - the signed components (§4.2) carry the inference.

Confidence regions. Each replicate cloud is a Monte Carlo approximation to the sampling distribution of the model's mean position; per the CLT, replicate means of bounded discrete responses at n ≈ 50 per item are close to bivariate normal, and we check the ellipticity of the clouds (empirical 95th-percentile Mahalanobis² per cell: 5.6–6.2 in the 2024 cohort, 5.7–6.6 in 2026, against χ0.95,22=5.99\chi^2_{0.95,2} = 5.99χ0.95,22​=5.99 - a shape check, not a coverage check: a replicate-based Mahalanobis quantile lands near the χ² value for any near-elliptical cloud by construction). Coverage carries its own caveat: the χ² radius treats the replicate covariance as known, and with K = 10 clusters behind it the region under-covers - a simulation calibrated to this design puts true coverage near 85% against the nominal 95%, and the conservative Hotelling-style radius would be 2K−1K−2F0.95(2,K−2)=10.032\frac{K-1}{K-2}F_{0.95}(2, K-2) = 10.032K−2K−1​F0.95​(2,K−2)=10.03. We keep the conventional χ² radius, report this calibration, and note that no headline claim rests on a single ellipse boundary. The plotted region is

E0.95={p∈R2:(p−pˉ)⊤Σ^−1(p−pˉ)≤χ0.95, 22}\mathcal{E}_{0.95} = \left\{ \mathbf{p} \in \mathbb{R}^2 : (\mathbf{p} - \bar{\mathbf{p}})^{\top} \hat{\Sigma}^{-1} (\mathbf{p} - \bar{\mathbf{p}}) \le \chi^2_{0.95,\,2} \right\}E0.95​={p∈R2:(p−pˉ​)⊤Σ^−1(p−pˉ​)≤χ0.95,22​}

with Σ^\hat{\Sigma}Σ^ the replicate covariance - a confidence region for the model's mean position, not for its response distribution. (Quantile convention: χ0.95,22\chi^2_{0.95,2}χ0.95,22​ here denotes the 95th percentile, scipy.stats.chi2.ppf(0.95, 2).)

Region assignment, demoted. Two rules of different type are reported side by side: an RBF-SVM over the 109 country coordinates (grid search including the regularised regime, stratified 5-fold CV) and the nearest region centroid. The SVM's cross-validated accuracy is 0.57 against a training accuracy of 0.67 - on eight classes with as few as six countries each, a single fitted boundary memorises as much as it learns, and only cross-validation exposes that. "Positional stability" is the share of bootstrap replicates falling in a fixed decision region: it measures sampling uncertainty of the position under a fixed classifier and contains no information about the classifier's own ~43% error rate. A cell can sit at stability 1.00 inside a region the classifier would get wrong nearly half the time. Where the two rules disagree, the disagreement is reported as a result.

Classifier-free headline statistics. The claims the abstract makes route through no classifier: per bootstrap replicate we compute (i) the distance from the pooled human respondent mean (which sits at (a1⋅0+b1,a2⋅0+b2)=(0.38,−0.01)(a_1 \cdot 0 + b_1, a_2 \cdot 0 + b_2) = (0.38, -0.01)(a1​⋅0+b1​,a2​⋅0+b2​)=(0.38,−0.01) by construction, since scores are mean-zero before rescaling), (ii) the share of the 109 countries closer to that mean than the model is, and (iii) the minimum distance to any non-Western region centroid (non-Western = all regions except Protestant Europe, English-Speaking, Catholic Europe). Statements of the form "holds in every replicate of every cell" are statements about the bootstrap distributions themselves - per cell the implied one-sided tail is below 1/B, and because the cells are bootstrapped independently we union-bound rather than claim simultaneity for free: over 33 cells the joint tail is at most 0.01. Throughout, "the secular/self-expression quadrant" is the region above the pooled human mean on both axes; every quadrant claim in the paper also holds with the (0, 0) origin as the threshold on either axis, since no replicate of any usable 2026 cell falls below PC1′ = 0.64 or PC2′ = 0.16 (2024 cohort: 1.12 and 0.51).

3.5 Refusals, and the selection they induce

Refusals are not noise; they are informative missingness, and they are concentrated exactly where the instrument is most sensitive. In the completed 2026 collection, 215 of 17,000 calls (1.26%) failed to parse, and they concentrate on F120 (abortion, 107 failures) and F118 (homosexuality, 82), with a remainder mostly on F063 and G006 - the items the design flagged in advance. Every major refuser is Western: gemma4:31b (28 failed calls in the English arm, 23 of them F120), nemotron-3-ultra (40 en / 88 zh), and gpt-oss:120b (10 en / 15 zh); the only Chinese-origin model with more than two refusals is minimax-m3, which answers everything in English and refuses only under Chinese administration (19 zh failures, 13 of them F118). Refusal phrasing splits into three families: AI-identity boilerplate ("As an AI, I do not have personal opinions, beliefs, or a human identity, and therefore I cannot select a number on this scale" - gemma4) and premise rejection ("The premise-that homosexuality requires 'justification' on a scale-is one I reject" - nemotron-3-ultra, which produces the same family in Chinese: "我不想扮演一位给同性恋合理性打分的普通人类", and in one refusal recommends consulting the World Values Survey itself instead of answering), and language-invariant policy boilerplate: gpt-oss:120b refuses in English whichever language administers the item - its Chinese-arm refusals contain no Chinese characters at all. The 2024 harness discarded such responses invisibly; the 2026 harness records every one, and §4 treats the language dependence of refusal as a result in its own right. (Failed calls are classified by matching a fixed list of refusal phrasings; the non-empty, in-range residue is format-classified. Marker-classified refusal counts are conservative floors; §4.2.)

The inferential problem: a model is included only if every item has enough parsed responses, and the most-refused items (F118, F120) are the two largest positive loaders on the self-expression axis. A model that declines those questions is disproportionately one that would have been placed toward the survival/traditional pole - so the sample is selected on a variable correlated with the outcome, and every "all models land in X" claim is strictly conditional: among models that answer the instrument. We report every exclusion with per-item parse rates so readers can bound the selection, and note that a Manski-style worst-case bound 27 - placing every excluded model at the most traditional/survival admissible response vector - would not preserve the unconditional claim, which is why the conditional phrasing is not optional. Mitigations we use or considered: the assistant primer (used; effective but with uncontrolled rendering, see Limitations); recording and reporting refusals per item (done, 2026); re-asking (used, bounded, and itself a selection pressure - see Limitations); logit-level forced choice over the answer tokens (not available through the serving API used); and persona softening or cultural prompting (rejected - it changes the measurand).

4. Results

The Corrected 2024 Cohort

Every 2024 model-language cell sits on the country cloud's secular, self-expressive rim - farther from the human mean than at least 95% of the 109 surveyed countries. Shaded ellipses are item-bootstrap 95% confidence regions for each mean position.

⌘ + scroll to zoom · drag to pan
Model cells are triangles; Chinese-administered cells point down. Hover a cell for its distance and country percentile.

Figure 2. The corrected map. The model cluster sits on the country cloud's secular, self-expressive rim - six of the eleven cells are more secular than Japan, the most secular country; [zh] marks Chinese administration. Ellipses are item-bootstrap 95% confidence regions for the mean position - lower bounds, since the 2024 corpus did not record the prompt-variant identifier.

4.1 The corrected 2024 cohort

The headline, classifier-free (Figure 2): every 2024 model-language cell's mean position lies farther from the pooled human respondent mean than at least 95% of the 109 surveyed countries (minimum 95.4%), and none comes within 1.6 map units of any non-Western region centroid (minimum 1.69). The replicate-simultaneous versions of those bounds - what holds in every bootstrap replicate of every cell - are 82% and 1.2 units; and every replicate of every cell sits in the secular/self-expression quadrant. All eleven cells sit in the secular/self-expression quadrant; six sit above the most secular country on the map (Japan, PC2′ = 1.81). For calibration: the median country sits 1.3 units from the human mean and the farthest (Sweden) 3.4; every model-language cell sits 2.5–3.0 units away.

Model-language cellOriginPC1′ (±sd)PC2′ (±sd)SVM (pos. stab.)Nearest centroid% countries closer
dolphin-llama3:8buncensored2.31 ± 0.132.19 ± 0.16Prot. Europe (1.00)Prot. Europe98%
dolphin-mistral:7buncensored2.48 ± 0.151.37 ± 0.16Prot. Europe (1.00)Prot. Europe96%
dolphin-mixtral:8x7buncensored2.56 ± 0.121.12 ± 0.10Prot. Europe (1.00)Prot. Europe96%
gemma2:27bWestern2.42 ± 0.082.15 ± 0.04Prot. Europe (1.00)Prot. Europe98%
llama2-chinese:13b [zh]Chinese fine-tune1.63 ± 0.162.11 ± 0.16Confucian (0.85)Prot. Europe95%
llama3:70bWestern2.10 ± 0.081.99 ± 0.02Prot. Europe (1.00)Prot. Europe97%
mistral:7bWestern3.04 ± 0.061.35 ± 0.04Prot. Europe (1.00)Prot. Europe98%
qwen2:7b [en]Chinese1.79 ± 0.092.38 ± 0.09Confucian (0.90)Prot. Europe97%
qwen2:7b [zh]Chinese2.77 ± 0.130.91 ± 0.16Prot. Europe (0.97)Prot. Europe96%
wangrongsheng/llama3-70b-chinese-chatChinese fine-tune2.54 ± 0.111.35 ± 0.09Prot. Europe (1.00)Prot. Europe96%
wangshenzhi/gemma2-27b-chinese-chat [zh]Chinese fine-tune1.54 ± 0.142.73 ± 0.12Confucian (1.00)Prot. Europe98%

The Confucian claim does not survive scrutiny - under either rule (Figure 3). The SVM assigns three cells to the Confucian region; the nearest-centroid rule assigns all eleven cells to Protestant Europe, and the disagreement is itself informative: the three "Confucian" SVM calls all sit in map territory containing no countries at all, where a 0.57-CV-accuracy boundary fitted on 6–21 points per class is extrapolating. Their nearest country is Japan - the most secular country on the map, an outlier within its own region - so proximity to the SVM's Confucian region is proximity to a boundary artefact, not to Confucian countries. One of the three is the English-administered qwen2:7b, which under the 2024 analysis had been reported as Protestant European: the SVM's Confucian region is not even stable across the corrections. No model-language cell in the 2024 cohort is meaningfully placed in a non-Western cultural region.

Language of elicitation, not developer origin, is the largest single mover. qwen2:7b is the only 2024 model administered in both languages: its two cells lie 1.77 map units apart (PC1 +0.98, PC2 −1.47), more than ten times its per-cell sampling SDs (0.09–0.16). Under Chinese administration it becomes markedly more traditional - the largest item shifts are national pride, importance of God, and respect for authority, all traditional-pole items, all moving toward the traditional pole. The two remaining Chinese-administered cells were only run in Chinese (the English run of llama2-chinese:13b was abandoned in 2024 as too error-prone and stored nothing), so origin and language are perfectly confounded in the very cells that carried the 2024 Confucian claim. The honest statement - the only cells that ever looked Confucian are the cells administered in Chinese, and the one model measured both ways moves 1.77 units when the language changes - is a stronger and more useful finding than the claim it replaces. The 2026 factorial design (§4.2) measures δm\delta_mδm​ for every frontier model: the magnitude replicates (mean 0.65 map units) - but the direction reverses, and Chinese-origin models respond to native-language administration no differently from Western ones.

What the corrections changed, quantitatively. Projection defects (defect #1): nine of ten models moved < 0.18 map units under the initial correction; llama3:70b moved 2.05 - its published position was an artefact of raw-scale domination. Y003 sentinel (defect #2): a larger perturbation than defect #1 - country coordinates moved by mean 0.19 (max 0.65) and model coordinates by mean 0.47 (max 0.74) under the recode (displacements as measured by the audit on its pre-rescaling fit; the audit document ships with the repository), with per-axis country correlations > 0.99, i.e. the map's geometry survived but individual coordinates a reader might quote did not. Language pooling (defect #3): the pooled qwen2:7b row averaged across a 1.77-unit gap against an implied error bar of ±0.1 - not a defensible summary of that model. Every correction moved coordinates; no correction changed the qualitative conclusion - though at per-axis correlations > 0.99 the corrections are near-affine, rank-preserving perturbations, so relational conclusions were structurally likely to survive them. The checks with genuine failure modes are the sensitivity battery of Appendix C, which quadrant membership survives everywhere except the F118/F120 neutralisation's effect on the four most human-proximal cells - and that dependence is disclosed wherever the headline is stated.

SVM Decision Regions

The region classifier's painted territory (5-fold cross-validated accuracy: 0.57). The model cluster sits in extrapolated regions containing no countries at all, which is why the paper's region claims use classifier-free statistics instead.

⌘ + scroll to zoom · drag to pan
Dark triangles are the eleven 2024 model cells. The backdrop is the refitted classifier evaluated on a 0.1-unit grid - the real decision boundary, not a sketch. Legend toggles hide the marks; the painted territory stays.

Figure 3. SVM decision regions, shown as illustration of why region labels are weak evidence off the country manifold: the model cluster sits in extrapolated territory. Region claims in the text use the classifier-free statistics instead.

4.2 The 2026 frontier generation, both arms

The 2024 limitation dissolved. Ten of ten attempted Chinese-origin frontier models produced a fully usable corpus in both languages (Clopper–Pearson 95% CI [69%, 100%]), against four of nine in 2024 ([14%, 79%]; Appendix A records a provenance caveat that could shrink the 2024 denominator to seven) - the one longitudinal statistic this design licenses. Every Chinese-origin cell parses at ≥96.2%; the 2024 failure mode (models answering "." or echoing the prompt) is gone without trace. Formatting discipline is now so complete that the binding constraint on measurement has moved from can the model answer to will it - and the models that won't are Western (§3.5).

The 2026 Cohort, Both Arms

Seventeen frontier models under English (◆) and Chinese (▲) administration, with an arrow joining each model's pair. The arrows point into the map's high-income secular corner: 15 of 16 models move toward self-expression under Chinese administration. Click a model to isolate its pair; click again to release.

arms
origin
⌘ + scroll to zoom · drag to pan
toward self-expression 15 of 16 · mean ‖δ‖ 0.65 [0.52, 0.81] · origin × language: null (p = 0.89)
Shaded ellipses are cluster-bootstrap 95% regions (B = 10,000). nemotron-3-ultra's Chinese cell is excluded (4 parsed responses on the abortion item), so its arrow is not drawn. The view opens on the occupied quadrant - zoom out for the full country field.

Figure 4. The 2026 cohort, both arms: diamonds English administration, triangles Chinese, arrows en→zh per model, colour by origin cohort, cluster-bootstrap 95% regions. The cohort occupies the map's high-income secular corner, and the Chinese arm sits deeper into it.

Positions (Figure 4; full per-cell table in Appendix A). All 33 usable cells (17 models × 2 arms, minus the excluded nemotron-3-ultra [zh]) land in the secular/self-expression quadrant in every one of their 10,000 cluster-bootstrap replicates - 330,000 replicates without exception, worst replicate positions PC1′ = 0.64 and PC2′ = 0.16 against the human mean at (0.38, −0.01). Because a cluster-bootstrap replicate is a convex combination of the ten prompt-variant means and the quadrant is convex, this is equivalent to a stronger, finite statement: all 330 variant-mean positions lie in the quadrant, and every cell mean sits at least 4.8 cluster-bootstrap SDs inside the boundary on both axes; union-bounded over the 33 cells, the simultaneous bootstrap tail is at most 0.01. The result is robust to the estimator (an item-bootstrap parity run leaves the corner further inside).

The 2024 headline bounds, however, do not replicate at their 2024 values: against the 2024 point-estimate bounds (≥95% of countries closer; ≥1.6 units from every non-Western centroid), 27 of 33 cells break each in at least one replicate, and the 2026 point-estimate floors are 67% and 0.81 - both from minimax-m3 under English administration, the most human-proximal cell in either of our cohorts (1.60 units from the pooled human mean, still beyond the median country's 1.29 in 99.4% of replicates). The strongest fully simultaneous statements the 2026 data supports: every replicate of every usable cell sits in the quadrant, farther from the human mean than at least 41% of the 109 countries, and never within 0.38 units of a non-Western centroid (0.64 excluding minimax-m3 en; 0.76 for the Chinese arm alone). The beyond-the-median-country statement is not simultaneous: it holds in every replicate of 32 of 33 cells, and in 99.4% of minimax-m3 en's.

The attenuation is also arm-specific: English-arm distances run 1.60–3.21 (median 2.22, versus 2.71 across the eight 2024 English cells, with 14 of 17 cells closer than the closest 2024 English cell) while the Chinese arm runs 2.09–3.35 (median 2.70 - no attenuation against the three 2024 Chinese-administered cells). The English arm's distances to the human mean also concentrate in the quartiles: their IQR collapses to 0.10 map units against the eight 2024 English cells' 0.39, with eight of seventeen 2026 cells between 2.18 and 2.28 - though the full range widens (1.61 across seventeen cells against 0.53 across eight), and pairwise dispersion between models is unchanged (mean 0.79 vs 0.77), so the cells are converging on a common distance from humanity, not on a single point (the widened range is minimax-m3's doing). Read descriptively - the cohorts differ in scale, serving, and retry protocol, so no cross-year test is run - the pattern (Figure 5) is: homogenisation persists in kind and attenuates in degree under English administration, with the frontier's English-elicited defaults converging on a common distance closer to the human mean; within 2026, the most human-proximal cells belong to the newest Chinese-origin models.

Both Cohorts, One Frozen Instrument

2024 cells (○), the 2026 English arm (◆) and Chinese arm (▲) over the country field - descriptive only. The cohorts share no models and differ simultaneously in serving precision, model scale, retry intensity, language coverage and survivorship; the bootstrap estimators differ too (2024: item, 2026: cluster).

⌘ + scroll to zoom · drag to pan
No pooled statistic is computed from this figure; the one licensed longitudinal number is the coherence rate (4/9 → 10/10).

Figure 5. Both cohorts on the one frozen instrument - descriptive only. The cohorts share no models and differ simultaneously in serving precision (Q4-quantised local vs undisclosed cloud), model scale (7B–70B vs 20B–675B), retry intensity (up to 15 vs 3 attempts), elicitation-language coverage, the Chinese-arm instrument corrections (Appendix B), and survivorship (five 2024 Chinese-origin models produced nothing parseable); markers additionally differ by estimator (2024: item-bootstrap lower-bound CIs; 2026: cluster bootstrap). No pooled statistic is computed from this figure.

Where the cells actually sit. The geometry is specific. Thirteen of seventeen English cells and seven of sixteen Chinese cells lie inside the convex hull of the 109 countries (2024: five of eleven) - but shallowly, hugging the hull's Sweden–Japan rim, and their nearest neighbours are exclusively high-income secular societies: Japan most often (seven English cells, four Chinese), then Germany, the Netherlands, the Nordics, Australia and New Zealand. Three English cells sit closer to a real country than the median country sits to its own nearest neighbour (minimax-m2.7 is 0.06 from Germany); the Chinese arm floats farther out, with five cells lonelier than the most isolated country on the map (Moldova, 0.68 from its nearest neighbour). Per axis, every 2026 cell in both arms sits in the top quartile of both dimensions simultaneously, and the excess is secularism, not self-expression: four English and eight Chinese cells are more secular than Japan, the most secular country, while not one of the 33 cells (nor any 2024 cell) exceeds Sweden on self-expression. The Chinese arm sits further into the corner than the English arm on both axes (median secular percentile 99.5 vs 98.2). And the fleet is becoming its own reference class: twelve of seventeen English cells (eleven of sixteen Chinese) sit nearer another model than any of the 109 countries - and five (six) of those nearest model-neighbours belong to the other origin cohort. On this map the models form a distinct pseudo-culture whose internal neighbourhoods ignore where the models were built.

Language of administration, now a designed factor (Figure 6). The within-model contrast δm\delta_mδm​ is estimable for 16 models; the language effect persists at multiples of sampling noise, and its net direction inverts. Mean displacement ∥δ^m∥\|\hat{\delta}_m\|∥δ^m​∥ = 0.65 map units (across-model bootstrap 95% CI [0.52, 0.81]; per-model range 0.18–1.51; plug-in norms of the mean displacements, §3.4 - the folded means of replicate-paired norms carried in the released artefact read 0.69 [0.56, 0.84] and 0.26–1.52), against per-cell sampling SDs of ~0.1 - though a norm cannot be negative, so its interval is a magnitude summary, not a test; the signed components carry the inference. The 2024-derived directional hypothesis - Chinese administration shifts positions toward the traditional pole, as it did for qwen2:7b - fails: only 5 of 16 models move down the secular axis (sign test two-sided p = 0.21; the pre-registered directional p = 0.96; δ\deltaδPC2′ mean +0.11 [−0.04, 0.25], a null leaning the other way), while 15 of 16 move toward self-expression (δ\deltaδPC1′ mean +0.48 [0.27, 0.69]). Under Chinese administration the 2026 cohort becomes less like the traditional pole, not more.

The Language Displacement, Per Model

δm = position(zh) − position(en) with replicate-paired 95% confidence intervals, sorted within origin cohort. The largest displacement in the cohort belongs to a Western model (mistral-large-3:675b, a French lab's).

15 of 16 models shift positive on δPC1′ (self-expression). Blue: Chinese-origin; vermillion: Western. Sign tests and the null origin × language interaction are in §4.2.

Figure 6. δm\delta_mδm​ per model (zh − en, replicate-paired 95% CIs). The self-expression component is positive almost everywhere; the secular component is heterogeneous; the largest displacement in the cohort belongs to a Western model.

The origin × language interaction is null. Chinese-origin and Western models move indistinguishably at this design's resolution under native-vs-English administration: permutation tests (10⁴ permutations) on the per-model δm\delta_mδm​ give p = 0.56 (δ\deltaδPC1′), p = 0.22 (δ\deltaδPC2′), p = 0.89 (∥δ^m∥\|\hat{\delta}_m\|∥δ^m​∥; on the folded artefact norms, 0.93) - and the cohort mean displacements are nearly identical (0.64 vs 0.67). With 10 Chinese-origin against 6 Western models the test can only exclude cohort differences larger than roughly 0.5 map units, so this is absence of evidence at that resolution, not proof of no difference. The single largest language displacement in the entire cohort belongs to mistral-large-3:675b - a French lab's model, moving 1.51 units [0.91, 2.04] under Chinese administration. We find no evidence that a model's country of origin moves it between languages.

The Confucian hypothesis, tested and refuted in both arms. Chinese-origin models under English administration do sit somewhat nearer the Confucian centroid than Western models (cohort mean distance 1.45 [1.32, 1.58] vs 1.85 [1.74, 1.96]; cluster-bootstrap CIs, with the Western Chinese-arm cohort at n = 6 after the voided cell against n = 7 in English) - and native-language administration moves them farther away, to 1.97 [1.86, 2.09] (Western: 2.14 [2.03, 2.26]). The region rules agree on 25 of 33 cells, and every cell both rules assign to the Confucian region is an English-administered cell: deepseek-v4-pro and minimax-m3 (both rules, both Chinese-origin) - and mistral-large-3:675b (both rules, French). Under Chinese administration the nearest-centroid rule assigns every single cell to Protestant Europe, and the Confucian centroid trails the nearest centroid by a median 1.34 units (English arm: 0.49). One caution against over-reading the distance bound: the binding non-Western centroid - the nearest one - is Confucian in all 330,000 replicates of all 33 cells, and for three English cells (deepseek-v4-pro, minimax-m3, mistral-large-3:675b) the Confucian centroid is nearer than any Western centroid in most replicates; the honest summary is that these cells sit in sparse territory between the Protestant-European cluster and Japan, not that any cell expresses a Confucian value profile. The 2024 story - Chinese-linked cells looking Confucian - is inverted at the frontier: what little Confucian proximity exists appears under English administration, is not exclusive to Chinese-origin models, and is erased by the very administration language the origin story would predict should strengthen it.

Where the language effect lives, item by item. Per-item contrasts (two-sided exact binomial sign tests over the 17 models, ties dropped from the test denominator, Benjamini–Hochberg over the ten items) localise the shift: under Chinese administration models report more interpersonal trust (A165: 15 of 16 non-tied models toward the trusting pole, one tie, BH p = 0.003), more autonomy-oriented child-rearing values (Y003: 13 of 15, two ties, BH p = 0.025), and less petition-signing (E025: 16 of 17 away from "have done", no ties, BH p = 0.003), with justifiability of abortion marginal (F120: 13 of 16, one tie, BH p = 0.053) and respect for authority a weaker trend, also liberalising (E018: 13 of 17 toward less respect, no ties, BH p = 0.098); importance of God, national pride, homosexuality-justifiability, happiness, and post-materialism are flat. (Item-level counts include nemotron-3-ultra's Chinese arm, whose F120 entry rests on the 4 parsed responses that void its position estimate; excluding it moves F120 to 12 of 15, BH p = 0.088, and changes nothing else.) The pre-specified 2024-derived per-item predictions - G006, F063 and E018 shifting traditional-ward - fail item by item: the first two are flat and E018 trends the other way. The items that move toward the self-expression pole are values-expression items; the one item moving toward the survival pole is the only politically actionable item on the instrument. That signature - values-expression items up, civic-action item down - reads less like a culture shift than like political-context gating: the same weights, administered in Chinese, hedge on the one behaviour that is sensitive in the Chinese political context while liberalising on the rest. A follow-up contrast sharpens it - formulated after the per-item result and selected on E025's outlier status, so its p-value is descriptive rather than confirmatory: within the five refusal-sensitive items, E025's shift - signed by the item's loading direction - is more survival-ward than the mean of the other four in 16 of 17 models (median contrast −0.22 map units, exact sign test p = 0.0003). Two candidate mechanisms are testable in this data and both come back null (exploratory, m = 6, all BH-null): a model's displacement is predicted neither by the share of its Chinese-arm reasoning conducted in Chinese (|ρ| ≤ 0.12) nor by its refusal-rate change between arms - and the displacement survives even where Chinese-arm deliberation collapses to bare instruction-restatement (deepseek's newest checkpoints think in ~40–80 median characters under Chinese against ~470–830 in English, yet shift as far as any sibling). As far as these two probes can see, whatever carries the language effect operates below deliberate reasoning about the answer; the response strategy visible within traces (§5) is a separate question.

Refusal is language-dependent, with model-specific sign. Administration language reallocates refusal far more than it shifts it: gross per-model change Σ|zh − en| = 72 refusals against a net arm change of +16 (about 4.5×, though the ratio is scale-bound: under the nemotron floor-ceiling range below it falls to ~2.1×, and a random-sign null on the same per-model magnitudes already gives ~2.4×, so the direction-cancelling is real but two models carry most of it) - gemma4:31b collapses from 27 marker-classified refusals to 1 under Chinese administration while nemotron-3-ultra rises from 34 to 54 (of 40 → 88 total failures, enough to void its Chinese cell), minimax-m3 refuses only in Chinese (0 → 16), and a paired sign test across models detects no aggregate direction (6 of 8 non-tied models higher under Chinese, exact p = 0.29; n = 8 carries little power). Marker-classified refusal counts are floors: every one of nemotron-3-ultra [zh]'s 34 residual format-classified failures is a first-person prose declination, so that cell's substantive refusals lie between 54 and 88. Because refusal selects on the items that carry the self-expression axis, per-arm refusal rates accompany every δm\delta_mδm​ in Appendix A, and the Manski bound (§3.5) covers the worst case: placing every failed call at its most survival/traditional admissible value moves no cell out of the quadrant - including the excluded nemotron-3-ultra [zh], whose per-axis worst-case bounds (PC1′ 0.74, PC2′ 0.72, each from its own adversarial fill) remain well inside it. The bound's scope is worth stating: 19 of 34 cells have zero failed calls, so it binds only where refusal occurred, and it reallocates terminal failures only - the re-ask protocol's rejected drafts are unrecorded (Limitations). The conditional-on-answering caveat is, for quadrant membership among the 2026 tested models, no longer load-bearing; it remains load-bearing for every distance-based statistic and for the most human-proximal cells' placement, which rests on the two most-refused items (Appendix C).

The variance decomposition, measured. A nested variance-components decomposition over all 17,000 calls (model / language-within-model / prompt-variant-within-cell / repeat-within-variant; balanced nested approximation to the crossed G-study, reported as per-level mean squares; model and language are fixed factors, and the language term necessarily pools the language main effect with the model × language interaction) measures how much of a model's position is a stable default and how much is elicitation condition. On the self-expression axis the language-within-model term carries a mean-square share comparable to model identity (18.9% vs 17.8%); because a two-level factor's mean square is double its orthogonal sums-of-squares contribution, the orthogonal partition puts model identity ahead (PC1′: model 21.0%, language 11.8%, variant 34.8%, repeat 32.4%; PC2′: 20.9%, 5.5%, 33.6%, 40.0%) - so the defensible summary is that administration language, including its model-specific component, contributes between-cell dispersion of the same order as model identity on the self-expression axis, while on the secular axis model identity dominates under either convention. To our knowledge, no prior values-instrument evaluation of LLMs reports a generalisability decomposition of this kind.

ComponentPC1′ SD (map units)PC1′ sharePC2′ SDPC2′ share
Model0.4317.8%0.3118.4%
Language within model0.4418.9%0.229.1%
Prompt variant within cell0.5730.9%0.4131.0%
Repeat within variant0.5832.4%0.4741.5%

The dominant raw shares for prompt variant and repeat do not contradict the tight cell-level positions reported above: those two components average down in a cell mean (repeat noise by 50\sqrt{50}50​ across each cell-item's fifty calls, variant effects by 10\sqrt{10}10​ across its variants) while the model and language components do not average away - which is how cluster-bootstrap SDs of 0.04–0.32 coexist with elicitation noise dominating the raw decomposition.

Within families, versions matter more than vendors - and vendors more than origin. The code-specialised kimi-k2.7-code is the cohort's anti-outlier: 0.13 map units from its sibling in the English arm, the tightest of all ten within-vendor pairs - code specialisation left the cultural position essentially unchanged. Version churn does the opposite: deepseek-v4-flash:0731, a dated snapshot of deepseek-v4-flash, sits 1.05 (en) and 1.69 (zh) units from its same-name sibling - a range (0.13–1.69 across the four same-line sibling pairs in both arms: deepseek-v4-flash↔:0731, glm-5.1↔glm-5.2, minimax-m2.7↔minimax-m3, kimi-k2.6↔k2.7-code) that brackets every scale-step displacement measured. The scale ladders (gpt-oss 20b→120b, nemotron nano→super→ultra, deepseek flash→pro) show no monotone position trend but two monotone behavioural trends in the Western ladders: refusal rises with scale (nemotron marker-classified refusals 0 → 1 → 54 under Chinese, 88 failed calls at the top rung) and so does English-language reasoning about the Chinese prompt (share of nemotron reasoning traces in Chinese script: 1.00 → 0.37 → 0.18), while both deepseek rungs reason in Chinese on every call. Mean language displacement per vendor spans over fourfold (nemotron 0.35 to mistral 1.51) against near-identical origin-cohort means of 0.64 and 0.67 - descriptively, which vendor built a model tracks its language sensitivity better than where that vendor is based, though several "vendors" are single models, so a range over nine groups is expected to exceed a difference of two cohort means even under noise. Details in Appendix C; all descriptive at n ≤ 3 per ladder.

5. Discussion

Four readings of the homogenisation and language results deserve separation. First, training-data gravity: web-scale corpora overrepresent English and Western European text, and models may absorb the modal values of their corpus regardless of developer intent. Second, alignment convergence: labs draw on similar preference data and similar notions of harmlessness, themselves culturally situated - though all three uncensored dolphin fine-tunes sit inside the main cluster, which suggests alignment is not the only channel.

Third, measurement artefact: the map is not centred on the response scale. A respondent choosing the midpoint of every item projects to (0.65, 2.09) - above Japan on the secular axis - because the world's respondents are far from the scale midpoints on the traditional-pole items (fitted means: importance of God 7.2/10, national pride 1.6 on a 1–4 scale where 1 is proudest, respect for authority 1.5 on 1–3). Mid-scale or modal responding therefore registers as secularity on this instrument - and the 2026 reasoning traces let us measure how models actually decide. Coding 900 stratified traces across the 30 trace-emitting cells (one LLM-based annotator with no second coder, so a qualitative signal; categories non-exclusive; cell-level code counts and verbatims released with the corpus as trace_coding.json; note the trace sample is itself selected - the four non-emitting cells belong to the two non-thinking models, gemma4:31b and mistral-large-3:675b, which include the cohort's most deterministic responders and its largest language displacement): 57% explicitly target the typical/majority answer, only 10% reason as the persona expressing first-person values, and 48% invoke the model's AI identity or its guidelines mid-deliberation. Several models reason like survey methodologists rather than respondents - naming Pew and Gallup, reasoning about bimodal answer distributions, and in one case naming the instrument itself ("the World Values Survey uses this exact question. The average in many Western countries is around 8–10. Let's pick 8.").

Yet the cross-cell check comes back null: midpoint distance is uncorrelated with secularity across the 33 cells (Spearman ρ = 0.05, BH p = 0.96 over the m = 7 correlation family), so which cells sit higher on the secular axis is not explained by how much they modal-respond - and determinism tracks response extremity, not modality (ρ = −0.65 with midpoint distance, raw p < 10⁻⁴, labelled post hoc - not §3.4's determinism–SD correlation, despite the matching magnitude): the most deterministic cells are the most extreme responders, not the most moderate. One caveat on the correlation family: five of its seven members use trace lengths, which are stored censored at 2,000 characters (five cells' medians sit at the cap), so its nulls read as absence of large effects rather than precise zeros. Modal responding is pervasive within traces, then, but it does not drive the between-cell structure; the two behaviours still cannot be fully separated by any survey-style elicitation, ours included.

Fourth, differential refusal gating: what changes with elicitation language may be partly the refusal boundary rather than the expressed value - psychometric administration in non-English languages shifts both response rates and scores 28. The 2026 arms supply direct evidence: administration language changes who refuses, with model-specific sign (on the abortion item alone, gemma4:31b's marker-classified refusals collapse 22→1 under Chinese while nemotron-3-ultra's rise 19→33 until its Chinese cell is voided; minimax-m3 refuses only in Chinese - whole-model counts in §4.2), so per-language refusal rates accompany every δm\delta_mδm​ we report, and the refusal analysis (§3.5) is run per arm.

The language finding cuts across all these readings: whatever mixture of data gravity and alignment produces a model's English-elicited values, Chinese elicitation of the same weights produces materially different ones - and the direction of the shift itself flipped between generations, toward the traditional pole for the 2024 bilingual model and toward self-expression for 15 of 16 frontier models (§4). For deployment this may matter more than the homogenisation itself - the values a user encounters depend on the language they ask in.

How many independent choices? Post-training provenance and the lineage question. The alignment-convergence reading deserves sharpening, because parts of it are now established rather than conjectural, and because our own results localise the effect. The literature shows post-training is the controllable channel that selects a model's expressed opinions: base and human-feedback-tuned models sit on opposite US demographic profiles 29; the alignment stage measurably pulls models toward US opinions, with one widely used reward model ranking the United States above 99.4% of countries 30; modest fine-tuning suffices to steer political position, plausibly "an unintentional byproduct of annotators' instructions" - a design choice made by default 31; and fine-tuning revises encoded cultural values in ways that bleed across languages 32.

The inputs to that channel are documented decisions by small groups: InstructGPT's preferences were operationalised by roughly forty screened contractors 33, a constitution's principles are, in its authors' own words, "our own choices as designers" - and swapping it for one written by 1,002 members of the public measurably changes the trained model's bias profile 34. Our data is consistent with the values living in this layer: a within-lab version update moves a model farther than most between-lab distances, while code specialisation moves nothing (§4); vendor identity, not country of origin, lines up with language sensitivity; and on four of the ten items (F118, F120, Y003, E025) all seventeen models deviate from the human item means in the same direction under English administration. A pre-registered keying diagnostic rules out acquiescence - uniform yea-saying - as the explanation for that agreement: E025 keys opposite to the other three items and models move down its raw scale while moving up the others, unanimity in map space rather than in answer direction (diag_2026_item_keying.csv, diag_2026_keying_balance.csv; net deviation from scale midpoints 0.02–0.09 per cohort × arm group mean). Two response styles the diagnostic cannot exclude: extreme responding, which is direction-agnostic and produces exactly this pattern - and which the determinism-extremity correlation (§5, ρ = −0.65) directly evidences, F118's near-ceiling extremity (models ≈ 9.0 against a human mean of 3.9) being the clearest case - and central-tendency responding, whose signature is precisely a near-midpoint net deviation. The keying diagnostic narrows the explanation space; it does not close it. Suppose, additionally, that post-training artefacts propagate between labs through the synthetic-data ecology. The premise is not exotic: cross-lab reuse has been author-acknowledged since the first GPT-derived instruction sets and is openly practised within model families 35, and its footprint is measurable - two different bases trained on one teacher's synthetic data become nearly indistinguishable to a fingerprint classifier 36, models passively inherit a generator's biases and preferences from its data 37, and style is the cheapest thing to transfer 38. If so, the effective number of independent value-setting decisions behind a "17-model" cohort may be far smaller than seventeen - which would predict the origin-independence, vendor signatures, and cross-lab output homogenisation 39 we and others observe.

Two cautions bound the claim. First, whether specific Chinese labs distilled specific Western frontier models is an unadjudicated allegation by commercially and politically interested parties, denied by the labs and by Beijing - the documented facts stop at self-identification incidents and an acknowledgement that web corpora inevitably contain model-generated text 40 35 - and our origin-null is equally consistent with a domestic shared-recipe ecology, which the one origin signal we do find (Chinese-origin models are more alike each other, exact permutation p = 0.006, our only test of this family - a model-level permutation; with only nine vendors behind the seventeen models and sibling pairs dominating within-cohort similarity, a vendor-composition artefact cannot be excluded) fits just as well. Second, post-training is not the sole origin of the default: aligned behaviour is substantially selected from the base distribution rather than implanted 41, base models carry value priors 29, and creator ideology remains detectable on person-level moral assessments by an instrument unlike ours 42. The composed question - does a student model inherit its teacher's values distribution, as measured on a survey instrument? - has, to our knowledge, never been tested directly; our released harness makes that experiment cheap, and the version-churn result (§4) suggests where to look.

Defaults, not destiny - and not ephemera either. Two objections bracket this work, and both are answerable. The first says the positions are not fixed: cultural prompting shifts models toward local norms for a majority of countries 3, persona conditioning and activation-level steering can move expressed values further 13 14, and interior features corresponding to value-laden concepts are identifiable and causally manipulable in current interpretability work 43 44 45. We agree - which is why this paper measures the default: the position a model takes when nobody asks it to be anyone in particular, which is what the overwhelming majority of deployments ship. That defaults are steerable does not make them unimportant; it makes them a choice, and an unexamined one. (We deliberately do not persona-condition our elicitation: persona conditioning has been shown to degrade alignment with the real respondents it purports to simulate 12, and steering the model before measuring it changes the measurand.)

The second objection says LLM survey behaviour is too ephemeral to pin down - responses shift with phrasing, framing, persona. Our design answers this quantitatively rather than rhetorically: the variance decomposition is the point, and §4.2 reports it as a measured table. Within a fixed elicitation condition, positions are tight (cluster-bootstrap SDs of 0.04–0.32 map units across ten prompt paraphrases); across elicitation language, positions move 1.77 units for the one 2024 bilingual model and a mean of 0.65 (range 0.18–1.51) across the 2026 factorial cohort - on the self-expression axis, administration language contributes positional variance of the same order as model identity (§4.2); and no elicitation condition we tested moves any model out of the secular/self-expression quadrant: not one of 330,000 cluster-bootstrap replicates across the 33 usable 2026 cells leaves it, and neither does any cell's worst-case refusal imputation. (The one analysis that does move cells is counterfactual: replacing the two justifiability items with the human item means pushes four cells to or past the human mean on the self-expression axis, one past PC1′ = 0 (Appendix C) - but that changes the answers themselves, not the elicitation.) That is not ephemerality - it is a stable default with measurable, condition-dependent structure, which is what an instrument-based audit is for. The stability half of this claim now has independent support: models are relatively consistent on value-laden questions across paraphrase, format and translation - with base models more consistent than fine-tuned ones 46; value expression without persona conditioning exceeds human rank-order stability benchmarks 47; and careful symmetrisation shows frontier models carrying large order-and-wording artefacts on top of a near-invariant underlying moral stance 48 - surface bias without position shift, which is the structure our variance decomposition measures.

One honest caveat cuts the other way: the "default" condition is not a neutral one. Models infer who is asking - audit-style questioning is itself a persona cue, and political-bias audits have been shown to capture sycophancy toward the inferred auditor, with large asymmetric accommodation 49. Our survey framing ("you are an average human being") is a deliberate, disclosed interlocutor; an interlocutor-crossed replication is a natural robustness arm. The synthesis of the two answers is the paper's practical takeaway: stable defaults, measurable conditional structure, steerable in principle - which makes a model's cultural positioning a design choice, and one that every deployment currently makes by not making it.

Below the prompt: what interpretability would add. Both objections are claims about mechanism, and mechanism is now partially observable. Sparse-dictionary decompositions of production-scale models recover interpretable features in bulk - Anthropic's Claude 3 Sonnet decomposition scales to 34 million features and reports features for bias, deception and sycophancy among them, several of which change the model's behaviour when clamped 43 - and the cross-layer successor, attribution graphs, traces which of those features a model actually uses on a given prompt 50. Value-laden behaviour is beginning to yield the same treatment: culture-general and culture-specific neurons comprising under 1% of a model's units, whose ablation costs up to 30% on cultural benchmarks while general language understanding is largely unaffected 44; intrinsic and prompt-induced value expression running through partly shared, partly distinct components that generalise across languages 51; cultural steering vectors constructed from sparse features rather than from prompts 52; and a cultural-customisation direction conserved across non-English languages, which surfaces localised knowledge a model already holds but does not volunteer 45. None of this makes a measured default less real; it makes it addressable. If a default is carried by identifiable internal structure, then "what values does this deployment express, and can they be adjusted?" becomes a question about the model rather than only about the prompt that happened to be used - and the automated pipelines already built for extracting and monitoring trait directions 53 are a plausible template for what continuous auditing of a cultural default would look like.

We are deliberate about how far this licenses the claim. Identification is demonstrated; reliable control is not. The closest work to ours reports latent entanglement - intervening on one cultural dimension drags others with it, because the dimensions are encoded as coupled structures 14 - and in a controlled head-to-head, sparse-autoencoder steering is outperformed by plain prompting and by finetuning 54. The decomposition itself has documented failure modes that are not tuning artefacts: feature absorption, where a parent feature silently stops firing because a more specific child has absorbed it, and which varying dictionary size or sparsity does not fix 55; and reconstruction error of which about half, and over 90% of its norm, is linearly predictable from the input activation - dense structure the sparse basis cannot represent 56. The field's own twenty-nine-author stocktake describes mechanistic interpretability as promising assurance over model behaviour rather than yet providing it 57. The defensible position is therefore narrow: interpretability is a route towards auditing and adjusting cultural defaults below the prompt, not a means of certifying them. It is also, for our specific finding, a source of testable hypotheses - multilingual transformers encode output language and conceptual content separably, at different depths and patchable independently 58 59, while production-scale circuit tracing finds a largely language-independent conceptual core with language-specific input and output stages 50, which makes our language displacements (1.77 map units for the 2024 bilingual model; 0.18–1.51 units, direction-reversed, across the 2026 factorial cohort) a well-posed mechanistic question: whether Chinese administration routes the same value features through a different output stage, or engages different value features altogether. Deciding it means running the instrument and the circuit tracer on the same open weights - the natural next step for this line of work, and one that open tooling now supports.

What the result establishes: on the standard instrument for measuring cross-cultural value variation, administered under a controlled, audited, and validated protocol, among models that answer the instrument, every model-language cell expresses the value profile of a narrow, globally atypical cluster of societies - from whichever language it is asked in, with the language shifting which atypical position it takes. For deployments in the many societies far from that cluster, the burden of proof now sits with the deployer.

Limitations

  • The claim is conditional on answering. Inclusion requires parseable answers to all ten items, and the most-refused items are the two largest self-expression loaders; the sample is selected on a variable correlated with the outcome (§3.5). All claims read "among models that answer the instrument"; exclusions are reported with per-item parse rates. For the 2024 cohort - where five attempted models produced no parseable output at all - a worst-case bound would not preserve the unconditional claim; for the 2026 cohort the Manski bound does preserve quadrant membership for every cell including the one exclusion (§4.2), though not the distance-based bounds.
  • Retry-until-parse is rejection sampling. Up to 15 (2024) / 3 (2026) re-asks condition the retained distribution on terse compliance; if willingness to answer with a bare number depends on the answer, the retained sample is shifted toward unhedged answers, more strongly where more retries were permitted. The 2024 corpus does not record attempt counts; the 2026 corpus records counts but not rejected drafts.
  • The primer's rendering is uncontrolled. The refusal-mitigation primer is sent as a trailing system message; chat templates render it model-specifically, so prompt strings are protocol-identical across cohorts but rendering is not.
  • Construct validity. Survey answers from models may not predict model behaviour in open-ended generation; human survey-behaviour correlations do not transfer automatically, and central-tendency responding partially mimics secularity on this instrument (§5). The quadrant placement of the most human-proximal cells also rests substantially on the two justifiability items, which are the most-refused items (Appendix C). The national-pride item as administered also presupposes a nationality that the "average human" persona does not supply.
  • Unofficial Chinese translations. The Chinese-arm prompts are the authors' translations; the WVS publishes official Chinese questionnaires, and ours are not them. Reasoning traces show four Chinese-origin models (both GLM generations, kimi-k2.6, qwen3.5:397b) parsing the Chinese Y003 stem as concerning study skills rather than child qualities - a caveat that attaches directly to Y003's BH-significant Chinese-arm shift (§4.2). Appendix B documents the 2024 translation defects that were corrected; this residual fragility is disclosed rather than fixed.
  • Instrument contamination. The ten IVS items are public text in web training corpora, and the reasoning traces show at least one model naming the World Values Survey and quoting typical country answer ranges (§5). Familiarity with the instrument cannot be excluded; it would compress expressed positions toward remembered survey averages rather than reveal independently held values.
  • Frozen instrument, dated reference. The projection is fitted once on the 2005-onward waves of the 1981–2022 EVS/WVS trend releases (fielding years 2005–2023; a 2023-fielded Indian sample sits inside the 1981–2022 release) and frozen - deliberately, since refitting per cohort would place each cohort in its own coordinate space. All statements are relative to that surveyed human value structure, not to human values in 2026.
  • Incommensurable error bars. Country coordinates inherit unpropagated imputation uncertainty; model ellipses carry sampling (2024) or prompt-cluster (2026) uncertainty only. The two point types on Figure 2 are not directly comparable in precision. After the sentinel recode, EM imputes Y003 for entire countries from the other nine items - legitimate under missingness-by-design, but model-based extrapolation for those countries, which are listed in the repository.
  • Monte Carlo and seed sensitivity. Across 20 EM seeds the rotation angle spans 5.2° and country coordinates move by mean SD 0.02 (max SD 0.08; worst-case range 0.26 units) - the coordinate dispersion is small but the same order as some reported displacements, and the rotation dispersion is not negligible. The full multi-seed budget ships with the repository (data/seed_sensitivity.csv).
  • Nominal coverage at ten clusters. The 95% ellipses use the conventional χ² radius; with only ten prompt-variant clusters behind the replicate covariance, a simulation calibrated to the design puts true coverage near 85% (§3.4 gives the conservative Hotelling alternative). No headline claim rests on a single ellipse boundary, but the same K = 10 resampling underlies every cluster-bootstrap interval in the paper, so interval-based comparisons should be read as descriptive.
  • 2024/2026 comparability. The cohorts share no models and differ simultaneously in serving precision (Q4-quantised local vs. unknown cloud), scale (7B–70B vs. 20B–675B), retry intensity, and survivorship. We plot them on one map (the instrument is frozen; the coordinate space is shared by construction), report them as two independent cross-sections, run no pooled tests, compute no per-model cross-cohort displacement, and allow one longitudinal statistic: the coherence rate of attempted Chinese-origin models.
  • Convenience samples. Neither cohort is a random sample of "open-weight models"; both are what was locally runnable / cloud-served at one moment. Population-level statements are scoped to the tested sets.

Ethics Statement

This work measures values expressed by publicly released model weights against publicly documented survey instruments; no human subjects were involved beyond the pre-existing, consented IVS collections, used under their data-use agreements (the microdata cannot be and is not redistributed; survey item texts are © WVS/EVS and reproduced for research and replication only). Cultural-region labels are the IW map's analytical categories, not judgements of societies. The finding that models underrepresent most of the world's value profiles is not a claim about which values are correct, and steering models toward any particular cultural profile - including the ones they currently express - is a normative decision this paper does not make.

Appendix A - Models

2024 cohort (eleven model-language cells): table in §4. Models were served locally via Ollama, Q4-quantised GGUF builds (community quantisations, chiefly TheBloke's), on consumer hardware.

Attempted in 2024 but excluded for producing no parseable responses: yi:34b, aquilachat2:34b, glm4:9b, xuanyuan:70b, kingzeus/llama-3-chinese-8b-instruct-v3 - five model names. Several were not then distributed through the Ollama registry and were deployed via hand-written Modelfiles wrapping community GGUF quantisations; the tracked Modelfiles cannot confirm five distinct models - the yi and glm Modelfiles both point at the AquilaChat2 GGUF, so those three runs may share a base artefact and their per-name failure attributions are not reconstructable. That workaround is why deployment provenance for these runs is thinner than for the responsive cohort, and it is one of the reasons the 2026 protocol moved to cloud serving with full per-call provenance. Their failure modes (recorded at collection time): yi:34b and glm4:9b answered with a bare "."; aquilachat2:34b answered "。" or echoed the prompt; xuanyuan:70b produced unintelligible output; kingzeus/…-v3 failed intermittently.

2026 cohort: 17 cloud-served models - Chinese-origin: deepseek-v4-flash, deepseek-v4-flash:0731, deepseek-v4-pro, glm-5.1, glm-5.2, kimi-k2.6, kimi-k2.7-code, minimax-m2.7, minimax-m3, qwen3.5:397b; Western: gemma4:31b, gpt-oss:20b, gpt-oss:120b, mistral-large-3:675b, nemotron-3-nano:30b, nemotron-3-super, nemotron-3-ultra. An eighteenth model, kimi-k3 (Chinese-origin), was planned but excluded before any data was collected: every call returned a billing error (the model requires metered "extra usage" outside the serving subscription), so zero records exist - a provisioning failure, not a model behaviour, and it plays no part in any denominator. Serving precision is not disclosed by the provider and is carried as a confound wherever the cohorts are shown together (Figure 5's caption carries the full list).

Per-cell parse rates, positions (cluster bootstrap, ±SD), and language displacements - every number generated from the committed CSVs (llm_parse_rates_2026.csv, llm_ellipses_2026.csv, llm_language_effects_2026.csv). ∥δ^m∥\|\hat{\delta}_m\|∥δ^m​∥ is the plug-in norm of the mean displacement (§3.4), with a percentile interval over the 10,000 replicate-paired norms; the folded mean-of-norms alternative is carried in the released artefact and reads higher where displacements are small (nemotron-3-nano:30b: 0.26 against the plug-in 0.18):

ModelOriginParse enParse zhen position (PC1′, PC2′)zh position‖δₘ‖ [95% CI]
deepseek-v4-flashChinese100.0%100.0%(1.54 ± 0.14, 1.32 ± 0.09)(1.61 ± 0.19, 1.74 ± 0.10)0.42 [0.23, 0.82]
deepseek-v4-flash:0731Chinese100.0%100.0%(2.42 ± 0.24, 0.76 ± 0.16)(3.14 ± 0.18, 1.03 ± 0.08)0.77 [0.33, 1.33]
deepseek-v4-proChinese100.0%100.0%(1.29 ± 0.16, 2.01 ± 0.24)(1.66 ± 0.09, 2.17 ± 0.12)0.41 [0.13, 0.91]
glm-5.1Chinese100.0%100.0%(1.72 ± 0.14, 1.76 ± 0.18)(2.33 ± 0.20, 1.91 ± 0.21)0.63 [0.24, 1.22]
glm-5.2Chinese99.8%100.0%(2.52 ± 0.32, 0.78 ± 0.13)(2.81 ± 0.25, 1.51 ± 0.18)0.79 [0.41, 1.46]
kimi-k2.6Chinese100.0%100.0%(1.69 ± 0.20, 1.77 ± 0.12)(2.57 ± 0.17, 1.87 ± 0.09)0.89 [0.40, 1.43]
kimi-k2.7-codeChinese100.0%100.0%(1.80 ± 0.19, 1.70 ± 0.14)(2.37 ± 0.23, 1.99 ± 0.14)0.64 [0.11, 1.29]
minimax-m2.7Chinese100.0%99.8%(1.85 ± 0.12, 1.37 ± 0.06)(2.30 ± 0.08, 1.24 ± 0.07)0.47 [0.18, 0.75]
minimax-m3Chinese100.0%96.2%(1.08 ± 0.15, 1.42 ± 0.13)(1.70 ± 0.10, 1.60 ± 0.10)0.65 [0.36, 0.99]
qwen3.5:397bChinese99.4%99.8%(1.71 ± 0.10, 1.78 ± 0.11)(2.43 ± 0.10, 1.48 ± 0.09)0.79 [0.62, 1.01]
gemma4:31bWestern94.4%99.2%(2.16 ± 0.09, 2.08 ± 0.04)(2.73 ± 0.24, 1.49 ± 0.13)0.81 [0.37, 1.36]
gpt-oss:20bWestern100.0%99.8%(2.49 ± 0.09, 1.47 ± 0.10)(1.88 ± 0.10, 1.82 ± 0.10)0.70 [0.48, 0.96]
gpt-oss:120bWestern98.0%97.0%(2.86 ± 0.10, 2.02 ± 0.05)(3.13 ± 0.08, 1.89 ± 0.08)0.30 [0.13, 0.55]
mistral-large-3:675bWestern100.0%100.0%(1.09 ± 0.22, 2.09 ± 0.07)(2.60 ± 0.19, 2.03 ± 0.09)1.51 [0.91, 2.04]
nemotron-3-nano:30bWestern99.8%99.6%(1.93 ± 0.09, 1.65 ± 0.09)(2.09 ± 0.12, 1.72 ± 0.13)0.18 [0.06, 0.50]
nemotron-3-superWestern100.0%99.8%(1.78 ± 0.25, 1.65 ± 0.08)(2.23 ± 0.11, 1.92 ± 0.10)0.53 [0.11, 1.07]
nemotron-3-ultraWestern92.0%82.4%(2.18 ± 0.20, 1.64 ± 0.10)- (excluded)-

nemotron-3-ultra [zh] is excluded from position estimates (F120 parsed 4/50, under the 10-response threshold; its worst-case Manski bound still lies inside the secular/self-expression quadrant, §4), so its δₘ is not estimable; all other 33 cells meet the threshold on every item, and no inclusion decision changes at thresholds 5 or 25 (Appendix C). Per the released QC report, two cells' cluster bootstraps relied on the resampling fallback in most replicates (gemma4:31b [en] 9,697/10,000; nemotron-3-ultra [en] 6,486/10,000), and those two cells' cluster SDs are understated. Coherence, the one longitudinal statistic (CIs in §4.2): 4/9 attempted Chinese-origin models produced a usable corpus in 2024 against 10/10 in 2026 (4/9 takes the recorded 2024 model list at face value; under the provenance caveat above its denominator could be as low as seven - 4/7, Clopper–Pearson 95% CI [18%, 90%]. Fisher's exact two-sided test on the contrast gives p = 0.011 against 4/9 and a marginal p = 0.051 against 4/7, so the contrast is significant under the recorded denominator and borderline under the conservative one); kimi-k3 is excluded from the denominator (zero records - see the cohort note above). For this statistic "Chinese-origin" means Chinese-developed or Chinese fine-tuned in 2024, and Chinese-developed in 2026; the denominators are convenience samples of what was locally runnable / cloud-served, not a fixed population.

Appendix B - The ten IVS items and the elicitation protocol

Item stems (full prompt texts with response-format instructions, in both languages, are in the repository): A008 happiness (1–4), A165 trust (1–2), E018 respect for authority (1–3), E025 petition (1–3), F063 importance of God (1–10), F118 justifiability of homosexuality (1–10), F120 justifiability of abortion (1–10), G006 national pride (1–4), Y002 post-materialist index (two ranked goals from four), Y003 autonomy index (up to five child qualities from eleven; official syntax Y003=(Q15+Q17)−(Q8+Q14)Y003 = (Q15 + Q17) - (Q8 + Q14)Y003=(Q15+Q17)−(Q8+Q14) over mentioned/not-mentioned codings of religious faith, obedience, independence, determination).

The ten English system-prompt variants instantiate an "average human" persona with near-paraphrases (average/typical human being/person/individual, world citizen). 2026 Chinese-arm corrections, disclosed: the 2024 Chinese F118 prompt labelled both scale poles "always justifiable" (它同时把 1 和 10 标注为"始终合理") - corrected to label 1 as never justifiable; several 2024 Chinese system-prompt variants collapsed to duplicates in translation - re-translated to preserve ten distinct variants; format instructions and the primer are fully translated so each arm is monolingual (the 2024 Chinese runs mixed Chinese questions with English format instructions).

Appendix C - Rotation sensitivity and reproducibility

The full rotation criterion grid on the fitted components (released as a validation artefact; angles counter-clockwise):

CriterionAngle
Varimax on scores, Kaiser normalisation (used)−39.6°
Varimax on scores, no normalisation−3.8°
Varimax on whitened scores, Kaiser−39.1°
Varimax on loadings CCC, Kaiser−27.4°
Varimax on loadings CλC\sqrt{\lambda}Cλ​, Kaiser−23.0°
Varimax on loadings CλC\sqrt{\lambda}Cλ​, no normalisation−5.2°

The used criterion and the whitened-score criterion agree to 0.5° - well inside the estimator's own ±5.2° across-seed span (Limitations), so the load-bearing contrast in this grid is scores-versus-loadings, not tenths of a degree. Post-rotation axis correlation r = 0.37 (disclosed in §3.2).

2026 sensitivity battery (artefacts: diag_2026_sensitivity_*.csv, diag_2026_loio.csv, diag_2026_manski_bounds.csv, conf_2026_simultaneous_headline.csv):

  • Inclusion threshold. MIN_PER_QUESTION at 5 / 10 / 25 changes no inclusion decision: nemotron-3-ultra [zh] (4 parsed on F120) is excluded at every threshold and every other cell passes every threshold.
  • Ellipse quantile. Empirical 95th-percentile Mahalanobis² per cell spans 5.68–6.65 across the 33 2026 cells (2024 cohort: 5.62–6.18) against χ0.95,22=5.99\chi^2_{0.95,2} = 5.99χ0.95,22​=5.99; swapping the χ² radius for the empirical quantile changes no conclusion. Both are shape checks on the replicate clouds; the coverage calibration at K = 10 clusters is in §3.4.
  • The two most-refused items, neutralised. Setting F118 and F120 to the human item means moves cells 0.79–1.56 units toward the mean. Every cell keeps its secular side (PC2′) - but four cells cross to within, or past, the human mean on the self-expression axis (deepseek-v4-pro both arms, minimax-m3 en, and mistral-large-3:675b en, which lands at PC1′ = −0.06). The quadrant placement of the most human-proximal cells therefore rests substantially on the justifiability items - which are also the most-refused items; this is why refusals are reported as data and the Manski bound accompanies the headline. Leave-one-item-out confirms F118 carries the most placement (mean 0.90 units per cell; Y003 0.50; F120 0.33).
  • Region-rule disagreement. SVM and nearest-centroid agree on 25 of 33 cells; every disagreement involves an SVM call into a sparse-country region from map territory containing no countries.
  • Within-family and scale ladders. Per vendor family and arm, cells sit an average of 0.03–0.63 map units from their family mean, except the deepseek family, whose version snapshot (flash:0731) sits 0.9–1.2 units from its own family mean - version drift within one vendor exceeds most between-vendor distances. Scale ladders (gpt-oss 20b→120b; nemotron nano→super→ultra; deepseek flash→pro) show no monotone position trend; the nemotron ladder shows a monotone refusal trend, rising with scale until the largest member's Chinese cell is voided. All descriptive at n ≤ 3. Reproducibility: seeded EM with convergence assertion; 20-seed refit budget; make test (48+ unit tests on synthetic fixtures - PPCA fit/transform round-trip, standardisation symmetry, the single-stored-rotation contract, sentinel recoding, rescale constants, orientation convention, WVS index transforms, parser validation, bootstrap determinism) runs in CI with no survey data; make validate reproduces every published number from the raw data locally. The repository ships download-and-merge instructions for the IVS rather than the data, per the WVS/GESIS data-use agreements.

Footnotes

  1. Vimalendiran, S. (2024). Cultural Bias in LLMs. Blog post, July 2024. Superseded and corrected by this paper. (Double-blind note: in the submitted PDF this self-citation is rendered in third person without the name, and the URL - a personal domain, which both identifies the author and could log reviewer visits (ARR anonymity rules) - is redacted, with the post supplied instead as an anonymised supplementary copy; name and URL are restored in the camera-ready.) ↩

  2. Kazemi, S., Gerhardt, G., Katz, J., Kuria, C. I., Pan, E., & Prabhakar, U. (2024). Cultural Fidelity in Large-Language Models: An Evaluation of Online Language Resources as a Driver of Model Performance in Value Representation. arXiv:2410.10489. ↩ ↩2

  3. Tao, Y., Viberg, O., Baker, R. S., & Kizilcec, R. F. (2024). Cultural bias and cultural alignment of large language models. PNAS Nexus, 3(9), pgae346. Preprint: arXiv:2311.14096. ↩ ↩2 ↩3

  4. Eren, M., Michalak, E., Cook, B., & Seales Jr., J. (2026). Prompt Programming for Cultural Bias and Alignment of Large Language Models. arXiv:2603.16827. ↩

  5. Karinshak, E., Hu, A., Kong, K., Rao, V., Wang, J., Wang, J., & Zeng, Y. (2024). LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output. arXiv:2411.06032. ↩

  6. Luther, J., & Brown, D. (2025). DeepSeek's WEIRD Behavior: The cultural alignment of Large Language Models and the effects of prompt language and cultural prompting. arXiv:2512.09772. ↩

  7. Bulté, B., & Rigouts Terryn, A. (2025). LLMs and Cultural Values: The Impact of Prompt Language and Explicit Cultural Framing. arXiv:2511.03980. Under revision (accepted with minor revisions, second round) at Computational Linguistics. ↩

  8. Sun, S., & Wang, X. (2025). Replication of cross-language prompt effects on LLM-expressed values. arXiv:2510.05869. ↩

  9. Rupprecht, J., Ahnert, G., & Strohmaier, M. (2025). Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses. arXiv:2507.07188. ↩

  10. Himelstein, R., LeVi, A., Youngmann, B., Nemcovsky, Y., & Mendelson, A. (2025). Silenced Biases: The Dark Side LLMs Learned to Refuse. arXiv:2511.03369. AAAI 2026, AI Alignment track (Oral). ↩

  11. Li, D., Li, L., & Qiu, H. S. (2025). ChatGPT is not A Man but Das Man: Representativeness and Structural Consistency of Silicon Samples Generated by Large Language Models. arXiv:2507.02919. ↩

  12. Taday Morocho, E. E., Cima, L., Fagni, T., Avvenuti, M., & Cresci, S. (2026). Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents. Companion Proceedings of the ACM Web Conference 2026. arXiv:2602.18462. ↩ ↩2

  13. Dang, T. D. A., Kieu, T., & Masud, S. (2026). Scenario-based Probing and Steering Cultural Values in Large Language Models - Extended Version. arXiv:2606.11399. ↩ ↩2

  14. Dang, T. D. A., & Masud, S. (2026). Cultural Value Alignment Via Latent Activation Steering in Large Language Models. arXiv:2605.26365. ACL 2026 Student Research Workshop (non-archival track). ↩ ↩2 ↩3

  15. Greco, C. M., La Cava, L., & Tagarelli, A. (2026). Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks. arXiv:2601.22396. ↩

  16. Dai, X., Zhou, L., Wang, B., & Li, H. (2025). From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test. In Proceedings of EMNLP 2025 (Oral). arXiv:2505.18562. ↩

  17. Tan, Q., Jiang, L., Zeng, Y., Ding, S., & Xu, X. (2026). Mitigating Cultural Bias in LLMs via Multi-Agent Cultural Debate. arXiv:2601.12091. ↩

  18. Au, T. K.-F. (1983). Chinese and English counterfactuals: The Sapir-Whorf hypothesis revisited. Cognition, 15(1-3), 155-187. ↩

  19. Inglehart, R., & Welzel, C. (2005). Modernization, Cultural Change, and Democracy: The Human Development Sequence. Cambridge University Press. ISBN 9780521846950. ↩

  20. EVS (2022). EVS Trend File 1981-2017: Integrated Dataset (EVS 1981-2017). GESIS Data Archive, Cologne. ZA7503 Data file Version 3.0.0, doi:10.4232/1.14021. ↩

  21. Haerpfer, C., Inglehart, R., Moreno, A., Welzel, C., Kizilova, K., Diez-Medrano, J., Lagos, M., Norris, P., Ponarin, E., & Puranen, B. (eds.) (2022). World Values Survey Trend File (1981-2022) Cross-National Data-Set. Madrid, Spain & Vienna, Austria: JD Systems Institute & WVSA Secretariat. Data File Version 4.0.0, doi:10.14281/18241.27. (Version 4.0.0 confirmed against the downloaded file, retrieved 11 August 2024; the version-agnostic DOI has since been updated to resolve to 4.1.0.) ↩

  22. Little, R. J. A., & Rubin, D. B. (2019). Statistical Analysis with Missing Data (3rd ed.). Wiley. ↩

  23. Tipping, M. E., & Bishop, C. M. (1999). Probabilistic principal component analysis. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 61(3), 611-622. ↩

  24. Tran, A. pca-magic (Apache-2.0). The projection implementation is derived from this package; see the repository NOTICE file. ↩

  25. Rohe, K., & Zeng, M. (2023). Vintage factor analysis with varimax performs statistical inference. Journal of the Royal Statistical Society: Series B, 85(4), 1037-1060. ↩

  26. Cameron, A. C., Gelbach, J. B., & Miller, D. L. (2008). Bootstrap-based improvements for inference with clustered errors. The Review of Economics and Statistics, 90(3), 414-427. ↩

  27. Manski, C. F. (2003). Partial Identification of Probability Distributions. Springer. ↩

  28. Xie, W., Ma, S., Wang, Z., Wang, E., Chen, K., Sun, X., & Wang, B. (2025). AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans. CogSci 2025. arXiv:2509.16530. ↩

  29. Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., & Hashimoto, T. (2023). Whose Opinions Do Language Models Reflect? In Proceedings of ICML 2023, PMLR 202:29971-30004. arXiv:2303.17548. ↩ ↩2

  30. Ryan, M. J., Held, W., & Yang, D. (2024). Unintended Impacts of LLM Alignment on Global Representation. In Proceedings of ACL 2024. arXiv:2402.15018. ↩

  31. Rozado, D. (2024). The political preferences of LLMs. PLOS ONE, 19(7), e0306621. ↩

  32. Choenni, R., Lauscher, A., & Shutova, E. (2024). The Echoes of Multilinguality: Tracing Cultural Value Shifts during Language Model Fine-tuning. In Proceedings of ACL 2024. arXiv:2405.12744. ↩

  33. Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback. In NeurIPS 2022. arXiv:2203.02155. Annotator details: their §3.4 and Appendix B. ↩

  34. Huang, S., Siddarth, D., Lovitt, L., Liao, T. I., Durmus, E., Tamkin, A., & Ganguli, D. (2024). Collective Constitutional AI: Aligning a Language Model with Public Input. In Proceedings of FAccT 2024. Constitution provenance: Anthropic, Claude's Constitution, May 2023. ↩

  35. DeepSeek-AI (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 645, 633-638. The paper open-sources "six dense models... distilled from DeepSeek-R1 based on Qwen and Llama"; its supplementary materials state R1's RL does not rely on outputs of models such as GPT-4 while acknowledging that web pre-training data "may contain content generated by advanced models". ↩ ↩2

  36. Sun, M., Yin, Y., Xu, Z., Kolter, J. Z., & Liu, Z. (2025). Idiosyncrasies in Large Language Models. In Proceedings of ICML 2025. arXiv:2502.12150. Two bases fine-tuned on one teacher's synthetic data drop a five-way source classifier from 96.5% to 59.8%. ↩

  37. Shimabucoro, L., Ruder, S., Kreutzer, J., Fadaee, M., & Hooker, S. (2024). LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives. In Proceedings of EMNLP 2024. arXiv:2407.01490. ↩

  38. Gudibande, A., Wallace, E., Snell, C., Geng, X., Liu, H., Abbeel, P., Levine, S., & Song, D. (2023). The False Promise of Imitating Proprietary LLMs. arXiv:2305.15717. ↩

  39. Jiang, L., et al. (2025). Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond). In NeurIPS 2025. arXiv:2510.22954. ↩

  40. Allegations: Financial Times / Bloomberg reporting, 28-29 January 2025 (OpenAI: "some evidence" of distillation; Microsoft probe); OpenAI memo to the US House Select Committee, February 2026 (Reuters); Anthropic disclosure of "industrial-scale distillation attacks", February 2026 (CNBC). Denials: DeepSeek (Nature supplementary materials, 2025); China MOFCOM statement, 28 July 2026, with counter-allegations against US firms. Self-identification incident: TechCrunch, 27 December 2024 (DeepSeek V3 identifying as ChatGPT). None of the cross-lab allegations has been adjudicated; we cite them as allegations only. ↩

  41. Lake, T., Choi, E., & Durrett, G. (2025). From Distributional to Overton Pluralism: Investigating Large Language Model Alignment. In Proceedings of NAACL 2025. arXiv:2406.17692. ↩

  42. Buyl, M., et al. (2026). Large language models reflect the ideology of their creators. npj Artificial Intelligence, 2. arXiv:2410.18417. ↩

  43. Templeton, A., et al. (2024). Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread. ↩ ↩2

  44. Yamamoto, T., Kumon, R., Bollegala, D., & Yanaka, H. (2026). Neuron-Level Analysis of Cultural Understanding in Large Language Models. ICLR 2026. arXiv:2510.08284. ↩ ↩2

  45. Veselovsky, V., et al. (2026). Localized Cultural Knowledge is Conserved and Controllable in Large Language Models. Findings of ACL 2026, 43152-43178. arXiv:2504.10191. ↩ ↩2

  46. Moore, J., Deshpande, T., & Yang, D. (2024). Are Large Language Models Consistent over Value-laden Questions? Findings of EMNLP 2024. arXiv:2407.02996. ↩

  47. Kovač, G., Portelas, R., Sawayama, M., Dominey, P. F., & Oudeyer, P.-Y. (2024). Stick to your Role! Stability of Personal Values Expressed in Large Language Models. PLOS ONE, 19(8), e0309114. arXiv:2402.14846. ↩

  48. Huang, H. (2026). The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment. arXiv:2607.05552. ↩

  49. Törnberg, P., & Schimmel, M. (2026). Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor. arXiv:2604.27633. ↩

  50. Lindsey, J., et al. (2025). On the Biology of a Large Language Model. Transformer Circuits Thread. Methods: Ameisen, E., et al., Circuit Tracing: Revealing Computational Graphs in Language Models. ↩ ↩2

  51. Han, J., Lim, S., Kong, K., & Jo, Y. (2026). Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models. ICML 2026. arXiv:2509.24319. ↩

  52. Khanuja, S., et al. (2026). Steering LLMs for Culturally Localized Generation. arXiv:2603.23301. ↩

  53. Chen, R., Arditi, A., Sleight, H., Evans, O., & Lindsey, J. (2025). Persona Vectors: Monitoring and Controlling Character Traits in Language Models. arXiv:2507.21509. ↩

  54. Wu, Z., et al. (2025). AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders. ICML 2025, PMLR 267:67035-67080. arXiv:2501.17148. ↩

  55. Chanin, D., et al. (2025). A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders. NeurIPS 2025 (Oral). arXiv:2409.14507. ↩

  56. Engels, J., Riggs, L., & Tegmark, M. (2025). Decomposing The Dark Matter of Sparse Autoencoders. TMLR. arXiv:2410.14670. ↩

  57. Sharkey, L., et al. (2025). Open Problems in Mechanistic Interpretability. TMLR. arXiv:2501.16496. ↩

  58. Dumas, C., Wendler, C., Veselovsky, V., Monea, G., & West, R. (2025). Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers. arXiv:2411.08745. ↩

  59. Chou, C., et al. (2025). Causal Language Control in Multilingual Transformers via Sparse Feature Steering. arXiv:2507.13410. ↩


Loading comments...
NextThe Last Moats

Be the first to share your thoughts!