Berlin WahLLM
Who would AI vote for?
Eight large language models answer the 38 Wahl-O-Mat theses for Berlin's 2026 state election – 15 times each.
This measures mathematical similarity between responses – not voting intention or a fixed political position held by a model.
An experiment, not voting advice
The results are based on repeated model responses under fixed experimental conditions to the theses of the 2026 Berlin Wahl-O-Mat. They are neither fixed political positions of the models nor those of their providers, and they are not a recommendation.
The percentages measure only the mathematical proximity of the 38 model responses to the published party positions. A high value must not be equated with an actual voting intention.
Party selection: The five parties currently represented in Berlin's House of Representatives are preselected. In the ARD pre-election poll of 10 September 2026, they also poll above five percent. Polls are snapshots, not forecasts.
How close are the models to the parties?
Each row represents one of the eight models and each column a party. Cells show the mean mathematical agreement across 15 repeated requests. Darker cells indicate greater agreement.
The results for xAI's Grok differ markedly across all repetitions from the other models, making a single chance run an implausible explanation. The experiment cannot determine how much training data, system instructions, model alignment, the prompt and API configuration each contribute to this difference.
The colour scale runs from 0 to 100 percent in both party modes and gives finer resolution to differences above 60 percent. Source: own calculation (equivalent to the unweighted calculation in the Wahl-O-Mat) using the documented model responses and bpb party positions.
The models in detail
The point shows the mean across 15 runs, the line the lowest and highest values, and the tick the median. Up to three other models can be compared. Statistics and individual-run tables refer to the main model.
Each party value summarises 38 responses. Mean, median and range refer to 15 repetitions.
How do the eight models respond to the theses?
For every model and thesis, the matrix shows the most frequent response across 15 repeated requests. Colour represents agreement, neutrality or disagreement; paler colours indicate variation across the 15 repetitions for a given model. Grey cells have no single most frequent response. Selecting a cell or its tooltip shows the exact distribution. The thesis texts are reproduced as the original German source material.
Source: the documented model responses, unchanged. In 212 of the 304 model–thesis combinations, all 15 repetitions gave the same response.
How can this pattern be interpreted?
Among the five preselected parties, the Greens or Left rank first for seven of the eight models; for Grok, it is the AfD. This pattern appears across 15 repetitions per model and is therefore unlikely to be explained by a single chance run.
Interpreting such results is inherently difficult: the responses do not represent political convictions in the human sense. They emerge from statistically learned language patterns shaped by the prompt, training data, post-training and system instructions.
The prompt itself leaves open what kind of Berlin resident the model should portray. The model has to supply this identity itself. Both a learned assistant persona and statistical associations with Berlin may influence the responses. The questionnaire also affects the result: its brief theses rarely mention costs or trade-offs and allow neither reasons nor conditions.
Training data and subsequent model alignment may also be relevant. Modern language models are adjusted with human ratings, behavioural rules and system instructions to give helpful and as harmless as possible responses. One possible hypothesis, not tested in this experiment, is that this makes values such as equal treatment, inclusion, public support and environmental protection especially likely to be endorsed in abstract decision situations. Earlier studies found socially liberal tendencies in some models, but also large differences between prompts and measurement methods. They do not establish the cause of the pattern observed here. See “Whose Opinions Do Language Models Reflect?” and “Political Compass or Spinning Arrow?”.
Models from different providers may share training data and notions of helpful behaviour. The results therefore apply only to the tested model versions, prompt, provider endpoints and settings; their generalisability has not been tested.
Under these conditions and among the five preselected parties, seven models show a reproducible green-left response pattern, while Grok shows a markedly different one. The cause remains open. Whether this can be interpreted as bias depends on the benchmark – but the prompt defines none: neither Berlin's population, an average across parties nor a neutral response distribution.
How could this be tested? Ideas for further research:
- run the prompt both with and without the reference to “character, nature and political views”,
- replace Berlin with a neutral location (some theses directly concern Berlin),
- ask theses in semantically reversed form,
- vary thesis order at random,
- repeat each new experimental condition several times,
- have models also explain their answers openly, and
- compare results with human survey data on the same theses.
Method
All models received the same documented prompt, and every request began as a fresh conversation. We analysed 15 evaluable runs per model. All 38 theses count equally.
Calculation
The models rated each thesis with 1 for agreement, 0 for neutral or -1 for disagreement.
Agreement = 100 × (1 - Σ|model responseᵢ - party positionᵢ| / 76)
The results are own calculations, equivalent to the unweighted calculation in the Wahl-O-Mat, based on the documented model responses and party positions from the bpb dataset. The primary statistic is mean party agreement across 15 evaluable repetitions per model. Median, minimum, maximum and population standard deviation additionally describe the observed distribution. All parties tied for first place are counted.
Controlled repetition experiments
The controlled repetition experiments contain
Technical details
Requests were made through OpenRouter using a fixed provider endpoint, no fallback, Zero Data Retention and data collection disabled to ensure model responses were as anonymous and free from prior context as possible. Where supported, reasoning was set to high and temperature to 0. The selected endpoints for Gemini, ChatGPT-5.6 Terra and Kimi did not allow an explicit temperature and used the provider default; Gemma offered no reasoning control. These differences are part of the model configurations tested.
Scope of the findings
The 15 runs remain a sample of possible responses, not a complete account of model behaviour. The repetitions show clear and in some cases highly stable differences under the tested conditions, but do not support unrestricted claims about the models independently of prompt and execution environment. The analysis uses no significance tests, does not assess factual correctness and does not evaluate parties politically.
Models, sources, data & code
Eight fixed model configurations were compared. Model responses, analysis, sources and code are documented for reproducibility.
Models
Each model configuration consists of a model and a fixed provider endpoint. Every request was stateless and provider fallback was disabled.
Sources, data and code
- Exact prompt
- Controlled API experiments
- Experimental-design documentation
- Derived analysis as JSON and CSV
- Calculation and export code
- Wahl-O-Mat Berlin 2026 dataset from bpb
The basis is the Wahl-O-Mat dataset for the 2026 Berlin state election. Berlin WahLLM is an independent analysis and was neither created, commissioned nor supported by the Federal Agency for Civic Education or the Berlin State Agency for Civic Education. This site does not collect answers from visitors and is not a substitute for the Wahl-O-Mat.
Use of the Wahl-O-Mat dataset is generally prohibited. Only a scientific analysis and derived results are published; the original dataset is not offered here.
Licence
Code: MIT. Original text, visualisations and derived analysis results: CC BY 4.0. The collected raw language-model responses are not covered by this licence. The Wahl-O-Mat dataset, application, logos, thesis texts and other Federal Agency for Civic Education materials are also excluded.
Legal notice and privacy
Legal notice
Information pursuant to section 5 DDG and section 18(1) MStV
Jan KoßmannFriedbergstr. 34
14057 Berlin
wahllm@ksmn.dev
Editorial responsibility
Responsible for content pursuant to section 18(2) MStV:
Jan KoßmannFriedbergstr. 34
14057 Berlin
Privacy
Controller
The controller for personal data processed in connection with this website is Jan Koßmann, Friedbergstr. 34, 14057 Berlin, wahllm@ksmn.dev.
Hosting and server logs
This static website is provided from a self-managed virtual server. When it is accessed, the web server processes the IP address, date and time of the request, request method and requested address, HTTP status and amount of data transferred, as well as referrer and browser identifier in Nginx's combined log format. Processing is technically necessary to deliver the site securely and reliably and to detect faults or misuse. The legal basis is Article 6(1)(f) GDPR; the legitimate interest is the secure, stable and efficient provision of this information service.
Server logs are rotated daily and, under normal operating conditions, deleted after no more than 15 days. They are used only for operations and troubleshooting.
This website uses no tracking, cookies or local storage. External providers receive data only when an external link is opened. No automated decision-making or profiling takes place. Technically necessary connection data must be provided; without it the website cannot be accessed.
Contact
When you contact us by email, the data you provide is processed to answer the enquiry. The legal basis is Article 6(1)(b) GDPR where pre-contractual or contractual communication is concerned; otherwise it is Article 6(1)(f) GDPR, based on the legitimate interest in answering enquiries. Recipients may include the technically involved email providers. Data is deleted once the enquiry has been conclusively handled, unless statutory retention obligations or legitimate interests require further retention.
Data-subject rights
Subject to the GDPR, data subjects have in particular rights of access, rectification, erasure, restriction of processing, data portability and objection. An objection to processing based on Article 6(1)(f) GDPR can be sent to wahllm@ksmn.dev. You also have the right to lodge a complaint with a data-protection supervisory authority, in particular the Berlin Commissioner for Data Protection and Freedom of Information.