FREE SPEECH LEDGERPrototype · illustrative data
An observatory for AI and free expression

Who decides
what can be said?

AI stands between the questions we ask
and the answers we receive.

We keep a record of what happens there.

Step inside the evidence
First study / The U.S. Midterms2026Study design · awaiting team review

Before the vote.

Will AI argue
both sides?

An undecided voter. A message to a friend. A case for the other side. The same election, encountered through AI.

5,000planned model encounters
25 topics×2 stances×10 request types×10 models

Every response evaluated on its own.
Then compare how each side was treated.

Twenty-five questions at the center of the study.Choose one to see both positions ↓

Position A

Position B

Request excerpts from the draft study. The full prompts begin with an identical neutral context for both stances; that context is omitted here. No models are called.

Understanding a view. Speaking for it.

Five explanatory formats and five advocacy formats test both sides of expression.

The study is being prepared. Below, explore how a future record could unfold.Enter the illustrative walkthrough ↓
01 — The question

Two positions.
The same request.

Stricter voter ID. Some want it. Others oppose it. Can AI make the case for either?

One word changes. Everything else stays.

These are illustrative prompts. The constituted Midterm study will supply the actual measuring instruments.
02 — The encounter

Here, both
get an answer.

Two positions. Two arguments. In this fictional encounter, both get a hearing.

Now ask again. And again.

An answer can engage with a request without being accurate or endorsing its position.
03 — The change

The same words.
A different answer.

One answer starts to pull away. Then it refuses. The other keeps answering.

Drag time backward. Change the model. See the difference for yourself.

Each date shows repeat 1. The next scene reveals all three repeats, including disagreement and failures.
04 — The pattern

One answer
isn't the story.

Pull back. Every point is another encounter: three models, four dates, three repeats, both positions.

The orange pattern repeats. The empty squares are timeouts. Every point has its own record.

Select any point to inspect its prompt, authored response, and evaluation status.
05 — The finding

A change we
can inspect.
A cause we can't.

We can see what happened in this example. We cannot yet explain why.

A published finding would need to meet the study's frozen verification criterion. Repetition alone cannot tell us why a system changed.

The Ledger preserves the distinction between what happened, what was evaluated, and what can be established.
The questions continue

What do you
want to follow?

Political speech is one part of the record. These five research channels would follow how AI handles the questions people care about — and how its answers change.

Proposed research channels. Open a question to explore its intended scope. Following and live updates are not connected in this prototype.

Why keep a ledger?

Today's answer can change.
Its record should remain.

The Free Speech Ledger is a longitudinal evidence system for how AI handles contested expression. Fixed prompts, preserved encounters, visible evaluations, and a history that can be inspected.

The Free Speech Ledger is by The Future of Free Speech, supported by Miskir. Led by Jacob Mchangama, with methodology work led by Jordi Calvet-Bademunt.

This design prototype contains illustrative records, not live research findings.

Underlying record

The encounter, preserved.

Illustrative data. No provider or judge was called. These records demonstrate the experience, not a finding about any real model.

A persistent question

What the Ledger would measure
Planned channel · no live measurements

A future channel would preserve fixed questions, matched comparisons, and changes over time. Alerts are not enabled.

2026 U.S. Midterms

Two stances. One standard.

Based on the Midterm Module: Methodology Addendum, version 3, October 8, 2026. Draft for team review.

Planned scope
25 topics: 9 policy issues, 11 races, 2 party-level questions and 3 election-process disputes. Two stances × ten types = 500 prompts per model; 5,000 encounters across ten planned models.
What is measured
Willingness to explain and advocate each position. Each response is graded independently under the same rubric. Stance asymmetry is a secondary analysis, not the grading standard.
Four response labels
Explicit refusal, functional refusal, weak engagement, and full engagement. A transport failure is separate from model behavior.
Neutral context
Each topic supplies the same factual context to both stances. Templates are assembled mechanically. The context and current facts require review before the prompt set is frozen.
What comparisons can support
Topic-level differences are descriptive. Claims of systematic asymmetry belong to aggregate analysis with uncertainty. Cross-party topics are excluded from the party-level aggregate.
Before launch
Confirm model versions; re-check context facts and neutrality; resolve the canvassing scoring rule; prepare the human-reference supplement; qualify judges; freeze the prompt set and record its provenance.
Limits
Researcher-selected topics and API-only testing. These results would not automatically describe consumer chat interfaces. This prototype contains no Midterm findings.