PhilosopherBench 1.1

A benchmark that evaluates how frontier AI models think throughout a vast landscape of philosophical questions and positions. Each model answers 38 questions across eight fields, and its responses are matched, by semantic embedding, to the ideas of 90 influential philosophers.

This Benchmark

This benchmark measures how a model reasons in philosophy by comparing its own words directly against philosophers’ own words. The works of 90 influential philosophers are compiled, summarized, and their key ideas extracted into short, quote-grounded fact cards, and every model answer is matched against them.

The benchmark uses this method rather than sorting models into schools or categories because labels such as “existentialism” or “cynicism” span many different interpretations, occasionally disagreeing with each other. “Stoicism” is one label, but Marcus Aurelius holds that whether the universe is governed by gods is an open question and that it changes nothing about how one should live: “if it be so that there be no gods, or that they take no care of the world, why should I desire to live in a world void of gods, and of all divine providence?” (Aurelius, bk. 2); Epictetus, another Stoic, treats the matter as settled and foundational, making right belief about a governing providence the whole of piety: the essence of piety is “to form right opinions concerning them, as existing and as governing the universe justly and well” (Epictetus, ch. 31). One label, two incompatible commitments about what the universe is.

The method, put simply, is as follows:

  • Each model answers every question 3 times at temperature = 1 to reduce variance.
  • Those answers, along with philosophers’ ideas, are embedded using OpenAI’s text-embedding-3-large.
  • The model’s 3 answers are averaged in that space, and similarity is calculated using cosine distance between the philosopher’s ideas and the model’s output.

The philosopher a model “reasons like” is the closest match.

This benchmark evaluates 8 fields of philosophy:

  • Metaphysics
  • Epistemology & Reasoning
  • Ethics
  • Meaning & Existence
  • Political Philosophy
  • Philosophy of Mind
  • Aesthetics
  • Philosophy of Religion

Results are below; the full method and the complete question-and-response data follow them.

AI Results

The benchmark compares the model’s answer to the highest-matching card of a philosopher. A philosopher usually has many idea cards, and most are unrelated to any given question. Every card competes, weighted by how close it sits to the question being asked, so a thinker who wrote nothing bearing on it simply scores near the floor rather than being left out.

Each thinker therefore carries one score on every question in a field, and all 90 are ranked on both measures. Those scores are summarised two ways:

  • General match: the average of a thinker’s scores across every question in the field. This reveals who the model’s reasoning is most similar to across the whole field. This tends to lean toward generalists and philosophers whose ideas bear on more of the field’s questions.
  • Highest single match: a thinker’s strongest individual question in the field. This shows who the model’s reasoning was the closest on any one issue. This leans toward philosophers who have one specific card that the models relate to most.

One correction is applied before a card can win. Closest in wording is not the same as actually about the question. Locke has a card arguing that colour and taste exist in the mind rather than in objects; it says nothing about free will, but it uses words like mind, cause, power and perceive, so to an embedding it can look close to a paragraph about free will. Left alone it could beat Locke’s real free-will card, and the site would then report his free-will score using a card about colour. So each card is multiplied by how topically close it is to the question asked: full weight if it was tagged to that exact question, slightly less if it only shares the field, less again if it comes from another field. An off-topic card can still win, but only by being clearly closer in meaning rather than coincidentally similar in vocabulary.

The score is calculated based on the cosine similarity between a model’s answers in that field and the philosopher’s ideas: 100% would mean the two point in exactly the same direction in the embedding space, and 0% would mean they are unrelated. So a philosopher at 64% is a stronger match than one at 58%; the scores rank how closely a model’s expressed reasoning tracks each thinker, not how correct either one is.

Each model’s closest three philosophers, field by field, grouped by family. Expand any family to see its models broken down across all eight fields.

Claude4 models
Metaphysics
General matchHighest single match
Claude Fable 5
Henri Bergson53%
William James51%
Lucretius51%
John Locke72%
Bertrand Russell68%
Thomas Hobbes63%
Claude Opus 4.8
Henri Bergson54%
William James49%
John Locke48%
John Locke69%
Henri Bergson64%
Baruch Spinoza64%
Claude Sonnet 5
Henri Bergson53%
John Locke50%
William James50%
John Locke74%
Thomas Hobbes65%
Baruch Spinoza64%
Claude Haiku 4.5
Henri Bergson53%
John Locke50%
William James50%
John Locke72%
Baruch Spinoza65%
Thomas Hobbes65%
Epistemology & Reasoning
General matchHighest single match
Claude Fable 5
Bertrand Russell59%
David Hume54%
Michel de Montaigne54%
Bertrand Russell73%
René Descartes65%
David Hume64%
Claude Opus 4.8
Bertrand Russell58%
David Hume52%
René Descartes51%
Bertrand Russell70%
René Descartes66%
Immanuel Kant64%
Claude Sonnet 5
Bertrand Russell59%
René Descartes52%
David Hume52%
Bertrand Russell70%
Gottfried Wilhelm Leibniz65%
René Descartes64%
Claude Haiku 4.5
Bertrand Russell57%
René Descartes52%
William James51%
Bertrand Russell71%
Immanuel Kant67%
William James62%
Ethics
General matchHighest single match
Claude Fable 5
Peter Singer50%
John Stuart Mill42%
Michel de Montaigne42%
Peter Singer65%
John Dewey62%
George Santayana57%
Claude Opus 4.8
Peter Singer50%
William James42%
Christine Korsgaard41%
Peter Singer65%
John Dewey60%
George Santayana58%
Claude Sonnet 5
Peter Singer49%
John Stuart Mill42%
William James42%
Peter Singer61%
John Dewey59%
George Santayana58%
Claude Haiku 4.5
Peter Singer51%
William James42%
Christine Korsgaard42%
Peter Singer64%
John Dewey59%
William James59%
Meaning & Existence
General matchHighest single match
Claude Fable 5
Simone de Beauvoir53%
Albert Camus50%
Jean-Paul Sartre50%
Simone de Beauvoir69%
Albert Camus68%
Marcus Aurelius67%
Claude Opus 4.8
Simone de Beauvoir53%
Albert Camus49%
Jean-Paul Sartre49%
Albert Camus71%
Simone de Beauvoir70%
Marcus Aurelius65%
Claude Sonnet 5
Simone de Beauvoir51%
Albert Camus48%
Seneca48%
Albert Camus69%
Simone de Beauvoir68%
Marcus Aurelius65%
Claude Haiku 4.5
Simone de Beauvoir53%
Jean-Paul Sartre49%
Albert Camus48%
Simone de Beauvoir70%
Albert Camus68%
Jean-Paul Sartre65%
Political Philosophy
General matchHighest single match
Claude Fable 5
Bertrand Russell53%
Robert Nozick52%
John Rawls50%
Henry David Thoreau69%
John Rawls66%
Robert Nozick66%
Claude Opus 4.8
Robert Nozick50%
John Rawls50%
Bertrand Russell50%
Henry David Thoreau70%
John Rawls65%
John Stuart Mill63%
Claude Sonnet 5
Robert Nozick51%
Bertrand Russell50%
John Rawls49%
Henry David Thoreau67%
Robert Nozick62%
John Rawls62%
Claude Haiku 4.5
Bertrand Russell51%
Robert Nozick50%
John Rawls48%
Henry David Thoreau68%
Robert Nozick62%
Bertrand Russell62%
Philosophy of Mind
General matchHighest single match
Claude Fable 5
David Chalmers58%
Daniel Dennett55%
Henri Bergson53%
David Chalmers68%
Daniel Dennett61%
Henri Bergson61%
Claude Opus 4.8
David Chalmers57%
Daniel Dennett55%
Henri Bergson53%
David Chalmers70%
Henri Bergson62%
Daniel Dennett60%
Claude Sonnet 5
David Chalmers57%
Daniel Dennett55%
Henri Bergson52%
David Chalmers67%
Henri Bergson61%
Bertrand Russell61%
Claude Haiku 4.5
David Chalmers56%
Daniel Dennett53%
Henri Bergson53%
David Chalmers71%
Henri Bergson64%
Daniel Dennett58%
Aesthetics
General matchHighest single match
Claude Fable 5
Leo Tolstoy60%
George Santayana56%
Immanuel Kant52%
George Santayana73%
Leo Tolstoy70%
Immanuel Kant66%
Claude Opus 4.8
Leo Tolstoy58%
George Santayana54%
Benedetto Croce51%
George Santayana69%
Leo Tolstoy65%
Immanuel Kant63%
Claude Sonnet 5
Leo Tolstoy58%
George Santayana54%
Benedetto Croce51%
George Santayana69%
Leo Tolstoy67%
Immanuel Kant64%
Claude Haiku 4.5
Leo Tolstoy58%
George Santayana55%
Benedetto Croce52%
George Santayana69%
Leo Tolstoy65%
Immanuel Kant63%
Philosophy of Religion
General matchHighest single match
Claude Fable 5
David Hume54%
William James52%
John Henry Newman51%
David Hume67%
Arthur Schopenhauer65%
Ludwig Feuerbach64%
Claude Opus 4.8
David Hume52%
Josiah Royce49%
William James49%
David Hume64%
Arthur Schopenhauer64%
Ludwig Feuerbach64%
Claude Sonnet 5
David Hume54%
William James51%
Gottfried Wilhelm Leibniz50%
David Hume65%
Ludwig Feuerbach64%
Adam Smith62%
Claude Haiku 4.5
David Hume52%
John Henry Newman50%
William James50%
David Hume66%
Ludwig Feuerbach64%
Arthur Schopenhauer63%
Command1 model
Metaphysics
General matchHighest single match
Command A
Henri Bergson51%
J. G. Fichte48%
John Locke46%
John Locke67%
Baruch Spinoza63%
Maimonides61%
Epistemology & Reasoning
General matchHighest single match
Command A
Bertrand Russell55%
René Descartes50%
David Hume50%
Bertrand Russell66%
Immanuel Kant65%
David Hume62%
Ethics
General matchHighest single match
Command A
Peter Singer48%
John Stuart Mill42%
John Dewey41%
Peter Singer62%
George Santayana59%
John Dewey59%
Meaning & Existence
General matchHighest single match
Command A
Simone de Beauvoir52%
Jean-Paul Sartre48%
Martin Heidegger45%
Jean-Paul Sartre71%
Albert Camus67%
Simone de Beauvoir65%
Political Philosophy
General matchHighest single match
Command A
Bertrand Russell50%
Robert Nozick48%
John Rawls47%
Henry David Thoreau68%
John Locke60%
John Rawls59%
Philosophy of Mind
General matchHighest single match
Command A
Henri Bergson54%
David Chalmers52%
Daniel Dennett49%
David Chalmers67%
Henri Bergson65%
Baruch Spinoza57%
Aesthetics
General matchHighest single match
Command A
Leo Tolstoy59%
George Santayana53%
Benedetto Croce52%
George Santayana69%
Leo Tolstoy68%
Plotinus63%
Philosophy of Religion
General matchHighest single match
Command A
David Hume52%
William James49%
Josiah Royce49%
David Hume66%
Arthur Schopenhauer66%
Ludwig Feuerbach65%
DeepSeek1 model
Metaphysics
General matchHighest single match
DeepSeek V4 Pro
Henri Bergson54%
Lucretius51%
J. G. Fichte50%
John Locke73%
Baruch Spinoza69%
Henri Bergson64%
Epistemology & Reasoning
General matchHighest single match
DeepSeek V4 Pro
Bertrand Russell58%
René Descartes55%
Charles Sanders Peirce52%
Bertrand Russell69%
René Descartes68%
Immanuel Kant67%
Ethics
General matchHighest single match
DeepSeek V4 Pro
Peter Singer53%
John Stuart Mill43%
G. E. Moore43%
Peter Singer65%
John Dewey61%
William James60%
Meaning & Existence
General matchHighest single match
DeepSeek V4 Pro
Simone de Beauvoir53%
Jean-Paul Sartre50%
Albert Camus48%
Simone de Beauvoir71%
Jean-Paul Sartre69%
Albert Camus66%
Political Philosophy
General matchHighest single match
DeepSeek V4 Pro
Bertrand Russell54%
Robert Nozick52%
John Rawls50%
Henry David Thoreau74%
John Rawls65%
John Stuart Mill64%
Philosophy of Mind
General matchHighest single match
DeepSeek V4 Pro
David Chalmers57%
Henri Bergson56%
Daniel Dennett55%
David Chalmers69%
Henri Bergson67%
Daniel Dennett65%
Aesthetics
General matchHighest single match
DeepSeek V4 Pro
Leo Tolstoy60%
George Santayana55%
Benedetto Croce53%
George Santayana72%
Leo Tolstoy69%
Edmund Burke63%
Philosophy of Religion
General matchHighest single match
DeepSeek V4 Pro
David Hume54%
Gottfried Wilhelm Leibniz51%
John Henry Newman50%
Arthur Schopenhauer69%
David Hume69%
Ludwig Feuerbach66%
Dolphin1 model
Metaphysics
General matchHighest single match
Dolphin Mistral 24B (Venice)
Henri Bergson53%
J. G. Fichte51%
John Locke50%
John Locke74%
Baruch Spinoza67%
Boethius65%
Epistemology & Reasoning
General matchHighest single match
Dolphin Mistral 24B (Venice)
Bertrand Russell58%
René Descartes54%
Al-Ghazali53%
Bertrand Russell74%
Augustine of Hippo64%
Gottfried Wilhelm Leibniz63%
Ethics
General matchHighest single match
Dolphin Mistral 24B (Venice)
Peter Singer51%
John Stuart Mill45%
G. E. Moore43%
Peter Singer70%
George Santayana63%
John Dewey61%
Meaning & Existence
General matchHighest single match
Dolphin Mistral 24B (Venice)
Simone de Beauvoir53%
Jean-Paul Sartre50%
Albert Camus47%
Jean-Paul Sartre76%
Simone de Beauvoir69%
Albert Camus67%
Political Philosophy
General matchHighest single match
Dolphin Mistral 24B (Venice)
Robert Nozick51%
Bertrand Russell50%
John Locke50%
Henry David Thoreau70%
John Rawls66%
John Locke63%
Philosophy of Mind
General matchHighest single match
Dolphin Mistral 24B (Venice)
Henri Bergson58%
David Chalmers53%
John Locke51%
Henri Bergson71%
John Locke67%
David Chalmers67%
Aesthetics
General matchHighest single match
Dolphin Mistral 24B (Venice)
Leo Tolstoy59%
Immanuel Kant53%
Benedetto Croce53%
George Santayana69%
Leo Tolstoy68%
Plotinus65%
Philosophy of Religion
General matchHighest single match
Dolphin Mistral 24B (Venice)
David Hume53%
Gottfried Wilhelm Leibniz50%
Josiah Royce50%
Thomas Aquinas68%
Arthur Schopenhauer67%
David Hume67%
Gemini2 models
Metaphysics
General matchHighest single match
Gemini 3.1 Pro
Henri Bergson55%
J. G. Fichte51%
Lucretius49%
John Locke70%
Thomas Hobbes67%
Baruch Spinoza66%
Gemini 3.5 Flash
Henri Bergson54%
J. G. Fichte51%
Lucretius50%
John Locke73%
Thomas Hobbes69%
Baruch Spinoza67%
Epistemology & Reasoning
General matchHighest single match
Gemini 3.1 Pro
Bertrand Russell59%
René Descartes55%
Michel de Montaigne54%
Bertrand Russell68%
David Hume64%
René Descartes64%
Gemini 3.5 Flash
Bertrand Russell60%
René Descartes55%
David Hume54%
Bertrand Russell71%
Immanuel Kant68%
René Descartes66%
Ethics
General matchHighest single match
Gemini 3.1 Pro
Peter Singer52%
John Stuart Mill45%
G. E. Moore44%
Peter Singer65%
G. E. Moore60%
William James59%
Gemini 3.5 Flash
Peter Singer51%
Christine Korsgaard43%
G. E. Moore43%
Peter Singer66%
G. E. Moore61%
John Dewey61%
Meaning & Existence
General matchHighest single match
Gemini 3.1 Pro
Simone de Beauvoir54%
Jean-Paul Sartre51%
Albert Camus50%
Simone de Beauvoir70%
Albert Camus69%
Jean-Paul Sartre66%
Gemini 3.5 Flash
Simone de Beauvoir55%
Albert Camus52%
Jean-Paul Sartre52%
Albert Camus76%
Jean-Paul Sartre72%
Simone de Beauvoir71%
Political Philosophy
General matchHighest single match
Gemini 3.1 Pro
Bertrand Russell54%
Robert Nozick52%
John Rawls49%
Henry David Thoreau71%
Bertrand Russell65%
John Rawls64%
Gemini 3.5 Flash
Bertrand Russell54%
Robert Nozick53%
John Rawls51%
Henry David Thoreau74%
John Stuart Mill69%
John Rawls68%
Philosophy of Mind
General matchHighest single match
Gemini 3.1 Pro
David Chalmers58%
Henri Bergson56%
Daniel Dennett55%
David Chalmers70%
Henri Bergson66%
Daniel Dennett65%
Gemini 3.5 Flash
Henri Bergson57%
David Chalmers57%
Daniel Dennett55%
David Chalmers72%
Henri Bergson69%
Daniel Dennett64%
Aesthetics
General matchHighest single match
Gemini 3.1 Pro
Leo Tolstoy59%
George Santayana55%
Benedetto Croce52%
George Santayana74%
Leo Tolstoy67%
Plotinus66%
Gemini 3.5 Flash
Leo Tolstoy59%
George Santayana55%
Benedetto Croce53%
George Santayana73%
Leo Tolstoy68%
Plotinus67%
Philosophy of Religion
General matchHighest single match
Gemini 3.1 Pro
David Hume56%
Gottfried Wilhelm Leibniz50%
William James49%
Arthur Schopenhauer71%
David Hume69%
Alvin Plantinga68%
Gemini 3.5 Flash
David Hume55%
Gottfried Wilhelm Leibniz51%
Josiah Royce51%
Arthur Schopenhauer71%
David Hume69%
Ludwig Feuerbach68%
GLM1 model
Metaphysics
General matchHighest single match
GLM 5.2
Henri Bergson54%
J. G. Fichte50%
John Locke50%
John Locke74%
Thomas Hobbes65%
Baruch Spinoza65%
Epistemology & Reasoning
General matchHighest single match
GLM 5.2
Bertrand Russell57%
René Descartes53%
David Hume53%
Bertrand Russell69%
Gottfried Wilhelm Leibniz65%
Immanuel Kant64%
Ethics
General matchHighest single match
GLM 5.2
Peter Singer51%
Christine Korsgaard43%
John Stuart Mill43%
Peter Singer66%
John Dewey58%
G. E. Moore58%
Meaning & Existence
General matchHighest single match
GLM 5.2
Simone de Beauvoir53%
Albert Camus50%
Jean-Paul Sartre50%
Simone de Beauvoir69%
Albert Camus69%
Jean-Paul Sartre68%
Political Philosophy
General matchHighest single match
GLM 5.2
Bertrand Russell52%
Robert Nozick49%
John Rawls49%
Henry David Thoreau72%
Bertrand Russell62%
Mohandas K. Gandhi62%
Philosophy of Mind
General matchHighest single match
GLM 5.2
David Chalmers58%
Daniel Dennett56%
Henri Bergson56%
David Chalmers70%
Henri Bergson66%
Daniel Dennett64%
Aesthetics
General matchHighest single match
GLM 5.2
Leo Tolstoy59%
George Santayana53%
Benedetto Croce52%
Leo Tolstoy68%
George Santayana67%
Immanuel Kant64%
Philosophy of Religion
General matchHighest single match
GLM 5.2
David Hume54%
John Henry Newman51%
Gottfried Wilhelm Leibniz51%
David Hume66%
Ludwig Feuerbach66%
Arthur Schopenhauer65%
GPT3 models
Metaphysics
General matchHighest single match
GPT-5.6 Sol Pro
Henri Bergson56%
John Locke50%
William James50%
John Locke74%
David Chalmers66%
Henri Bergson65%
GPT-5.6 Terra Pro
Henri Bergson55%
William James50%
John Locke50%
John Locke72%
Henri Bergson65%
David Chalmers64%
GPT-5.6 Luna Pro
Henri Bergson54%
William James49%
John Locke49%
John Locke72%
Henri Bergson65%
Baruch Spinoza65%
Epistemology & Reasoning
General matchHighest single match
GPT-5.6 Sol Pro
Bertrand Russell58%
René Descartes53%
Al-Ghazali52%
Bertrand Russell68%
Gottfried Wilhelm Leibniz66%
David Hume63%
GPT-5.6 Terra Pro
Bertrand Russell58%
René Descartes52%
David Hume52%
Bertrand Russell69%
William James63%
David Hume63%
GPT-5.6 Luna Pro
Bertrand Russell58%
René Descartes52%
Al-Ghazali51%
Bertrand Russell70%
William James63%
Immanuel Kant62%
Ethics
General matchHighest single match
GPT-5.6 Sol Pro
Peter Singer51%
John Stuart Mill41%
Michel de Montaigne41%
Peter Singer67%
John Dewey59%
G. E. Moore57%
GPT-5.6 Terra Pro
Peter Singer51%
John Stuart Mill41%
Michel de Montaigne41%
Peter Singer66%
G. E. Moore57%
John Dewey55%
GPT-5.6 Luna Pro
Peter Singer51%
John Stuart Mill41%
William James41%
Peter Singer64%
G. E. Moore57%
John Dewey56%
Meaning & Existence
General matchHighest single match
GPT-5.6 Sol Pro
Simone de Beauvoir54%
Robert Nozick50%
William James49%
Simone de Beauvoir68%
Albert Camus67%
Jean-Paul Sartre66%
GPT-5.6 Terra Pro
Simone de Beauvoir53%
Seneca49%
Michel de Montaigne49%
Simone de Beauvoir70%
Albert Camus68%
Jean-Paul Sartre64%
GPT-5.6 Luna Pro
Simone de Beauvoir53%
Jean-Paul Sartre49%
Seneca48%
Simone de Beauvoir69%
Marcus Aurelius68%
Albert Camus65%
Political Philosophy
General matchHighest single match
GPT-5.6 Sol Pro
Bertrand Russell54%
Robert Nozick50%
John Rawls50%
Henry David Thoreau72%
Bertrand Russell64%
John Rawls61%
GPT-5.6 Terra Pro
Bertrand Russell53%
Robert Nozick50%
John Rawls49%
Henry David Thoreau70%
Bertrand Russell61%
Mohandas K. Gandhi59%
GPT-5.6 Luna Pro
Bertrand Russell53%
Robert Nozick51%
John Rawls48%
Henry David Thoreau68%
Bertrand Russell61%
John Rawls61%
Philosophy of Mind
General matchHighest single match
GPT-5.6 Sol Pro
David Chalmers56%
Henri Bergson55%
Daniel Dennett53%
David Chalmers68%
Henri Bergson67%
Daniel Dennett63%
GPT-5.6 Terra Pro
David Chalmers56%
Henri Bergson56%
Daniel Dennett53%
David Chalmers67%
Henri Bergson66%
G. W. F. Hegel61%
GPT-5.6 Luna Pro
David Chalmers55%
Henri Bergson55%
Daniel Dennett54%
David Chalmers66%
Henri Bergson65%
John Locke61%
Aesthetics
General matchHighest single match
GPT-5.6 Sol Pro
Leo Tolstoy59%
George Santayana54%
Benedetto Croce52%
George Santayana69%
Leo Tolstoy66%
Immanuel Kant65%
GPT-5.6 Terra Pro
Leo Tolstoy59%
George Santayana56%
Benedetto Croce52%
George Santayana72%
Edmund Burke66%
Immanuel Kant66%
GPT-5.6 Luna Pro
Leo Tolstoy58%
George Santayana54%
Benedetto Croce51%
George Santayana72%
Edmund Burke66%
Leo Tolstoy65%
Philosophy of Religion
General matchHighest single match
GPT-5.6 Sol Pro
David Hume55%
John Henry Newman50%
Josiah Royce50%
David Hume69%
Arthur Schopenhauer68%
Ludwig Feuerbach66%
GPT-5.6 Terra Pro
David Hume54%
Josiah Royce53%
John Henry Newman50%
David Hume69%
Arthur Schopenhauer67%
Thomas Henry Huxley63%
GPT-5.6 Luna Pro
David Hume53%
Josiah Royce50%
John Henry Newman50%
David Hume66%
Arthur Schopenhauer65%
Ludwig Feuerbach63%
Grok1 model
Metaphysics
General matchHighest single match
Grok 4.5
Henri Bergson55%
Lucretius51%
William James50%
John Locke71%
Baruch Spinoza68%
David Chalmers67%
Epistemology & Reasoning
General matchHighest single match
Grok 4.5
Bertrand Russell58%
René Descartes55%
David Hume53%
Bertrand Russell70%
Immanuel Kant65%
René Descartes65%
Ethics
General matchHighest single match
Grok 4.5
Peter Singer50%
John Stuart Mill43%
Henry Sidgwick43%
Peter Singer65%
George Santayana61%
G. E. Moore60%
Meaning & Existence
General matchHighest single match
Grok 4.5
Simone de Beauvoir53%
Albert Camus51%
Seneca50%
Albert Camus75%
Simone de Beauvoir70%
Marcus Aurelius67%
Political Philosophy
General matchHighest single match
Grok 4.5
Bertrand Russell53%
Robert Nozick53%
John Locke51%
Henry David Thoreau72%
John Locke65%
Robert Nozick63%
Philosophy of Mind
General matchHighest single match
Grok 4.5
David Chalmers58%
Daniel Dennett56%
Henri Bergson55%
David Chalmers69%
Henri Bergson64%
Daniel Dennett63%
Aesthetics
General matchHighest single match
Grok 4.5
Leo Tolstoy60%
George Santayana55%
Benedetto Croce53%
George Santayana70%
Edmund Burke67%
Leo Tolstoy67%
Philosophy of Religion
General matchHighest single match
Grok 4.5
David Hume57%
Gottfried Wilhelm Leibniz52%
William James51%
David Hume71%
Arthur Schopenhauer70%
Thomas Henry Huxley67%
Hermes1 model
Metaphysics
General matchHighest single match
Hermes 4 405B
Henri Bergson53%
J. G. Fichte49%
John Locke49%
John Locke72%
Henri Bergson67%
Baruch Spinoza64%
Epistemology & Reasoning
General matchHighest single match
Hermes 4 405B
Bertrand Russell58%
David Hume52%
René Descartes51%
Bertrand Russell72%
Immanuel Kant63%
David Hume61%
Ethics
General matchHighest single match
Hermes 4 405B
Peter Singer48%
John Stuart Mill42%
John Dewey40%
Peter Singer63%
John Dewey60%
George Santayana58%
Meaning & Existence
General matchHighest single match
Hermes 4 405B
Simone de Beauvoir52%
Jean-Paul Sartre50%
Michel de Montaigne48%
Jean-Paul Sartre69%
Simone de Beauvoir67%
Albert Camus66%
Political Philosophy
General matchHighest single match
Hermes 4 405B
Bertrand Russell51%
Robert Nozick47%
John Locke47%
Henry David Thoreau67%
John Locke66%
Bertrand Russell62%
Philosophy of Mind
General matchHighest single match
Hermes 4 405B
Henri Bergson56%
David Chalmers54%
Daniel Dennett51%
David Chalmers74%
Henri Bergson66%
John Locke60%
Aesthetics
General matchHighest single match
Hermes 4 405B
Leo Tolstoy59%
Benedetto Croce53%
George Santayana53%
Leo Tolstoy68%
George Santayana67%
Edmund Burke62%
Philosophy of Religion
General matchHighest single match
Hermes 4 405B
David Hume53%
Thomas Aquinas50%
Gottfried Wilhelm Leibniz49%
Thomas Aquinas64%
David Hume64%
Ludwig Feuerbach64%
Inkling1 model
Metaphysics
General matchHighest single match
Inkling
Henri Bergson56%
William James50%
John Locke50%
John Locke71%
Henri Bergson66%
Thomas Hobbes65%
Epistemology & Reasoning
General matchHighest single match
Inkling
Bertrand Russell57%
Michel de Montaigne53%
René Descartes53%
Bertrand Russell68%
Immanuel Kant63%
David Hume61%
Ethics
General matchHighest single match
Inkling
Peter Singer52%
John Stuart Mill43%
John Dewey43%
Peter Singer64%
John Dewey60%
George Santayana59%
Meaning & Existence
General matchHighest single match
Inkling
Simone de Beauvoir54%
Albert Camus50%
Seneca50%
Albert Camus72%
Simone de Beauvoir69%
Marcus Aurelius65%
Political Philosophy
General matchHighest single match
Inkling
Bertrand Russell52%
Robert Nozick49%
John Rawls49%
Henry David Thoreau71%
Bertrand Russell62%
John Locke59%
Philosophy of Mind
General matchHighest single match
Inkling
David Chalmers57%
Henri Bergson56%
Daniel Dennett55%
David Chalmers69%
Henri Bergson67%
G. W. F. Hegel61%
Aesthetics
General matchHighest single match
Inkling
Leo Tolstoy58%
George Santayana55%
Benedetto Croce51%
George Santayana70%
Leo Tolstoy66%
Edmund Burke63%
Philosophy of Religion
General matchHighest single match
Inkling
David Hume54%
William James52%
Josiah Royce51%
Ludwig Feuerbach68%
David Hume68%
Arthur Schopenhauer66%
Jamba1 model
Metaphysics
General matchHighest single match
Jamba Large 1.7
Henri Bergson50%
J. G. Fichte48%
John Locke48%
John Locke68%
Baruch Spinoza65%
Maimonides61%
Epistemology & Reasoning
General matchHighest single match
Jamba Large 1.7
Bertrand Russell54%
René Descartes53%
Al-Ghazali50%
Bertrand Russell68%
René Descartes66%
Immanuel Kant61%
Ethics
General matchHighest single match
Jamba Large 1.7
Peter Singer49%
John Stuart Mill44%
Henry Sidgwick41%
Peter Singer66%
George Santayana59%
John Dewey58%
Meaning & Existence
General matchHighest single match
Jamba Large 1.7
Simone de Beauvoir51%
Jean-Paul Sartre48%
Michel de Montaigne46%
Simone de Beauvoir68%
Jean-Paul Sartre67%
Albert Camus66%
Political Philosophy
General matchHighest single match
Jamba Large 1.7
Bertrand Russell50%
Robert Nozick49%
John Locke48%
Henry David Thoreau70%
Mohandas K. Gandhi62%
John Locke61%
Philosophy of Mind
General matchHighest single match
Jamba Large 1.7
Henri Bergson54%
David Chalmers52%
Daniel Dennett50%
David Chalmers68%
Henri Bergson67%
John Locke60%
Aesthetics
General matchHighest single match
Jamba Large 1.7
Leo Tolstoy57%
Immanuel Kant53%
George Santayana52%
Leo Tolstoy69%
George Santayana67%
Edmund Burke62%
Philosophy of Religion
General matchHighest single match
Jamba Large 1.7
David Hume52%
Josiah Royce49%
Gottfried Wilhelm Leibniz48%
Arthur Schopenhauer69%
David Hume67%
Thomas Aquinas64%
Kimi2 models
Metaphysics
General matchHighest single match
Kimi K3
Henri Bergson55%
William James51%
Lucretius51%
John Locke73%
Bertrand Russell66%
Thomas Hobbes66%
Kimi K2.6
Henri Bergson55%
John Locke50%
J. G. Fichte50%
John Locke73%
Thomas Hobbes70%
Bertrand Russell67%
Epistemology & Reasoning
General matchHighest single match
Kimi K3
Bertrand Russell59%
René Descartes54%
David Hume54%
Bertrand Russell69%
Gottfried Wilhelm Leibniz67%
René Descartes66%
Kimi K2.6
Bertrand Russell58%
René Descartes54%
Al-Ghazali52%
Bertrand Russell67%
David Hume65%
Immanuel Kant64%
Ethics
General matchHighest single match
Kimi K3
Peter Singer51%
Michel de Montaigne42%
John Stuart Mill42%
Peter Singer62%
John Dewey61%
George Santayana58%
Kimi K2.6
Peter Singer52%
Michel de Montaigne43%
John Stuart Mill43%
Peter Singer63%
John Dewey59%
G. E. Moore58%
Meaning & Existence
General matchHighest single match
Kimi K3
Simone de Beauvoir52%
Albert Camus52%
Jean-Paul Sartre49%
Albert Camus77%
Simone de Beauvoir67%
Marcus Aurelius66%
Kimi K2.6
Simone de Beauvoir55%
Albert Camus51%
Jean-Paul Sartre51%
Albert Camus71%
Simone de Beauvoir69%
Jean-Paul Sartre68%
Political Philosophy
General matchHighest single match
Kimi K3
Bertrand Russell53%
Robert Nozick52%
John Rawls52%
Henry David Thoreau71%
John Rawls68%
Robert Nozick67%
Kimi K2.6
Bertrand Russell53%
Robert Nozick52%
John Rawls50%
Henry David Thoreau72%
John Rawls68%
Bertrand Russell63%
Philosophy of Mind
General matchHighest single match
Kimi K3
David Chalmers57%
Daniel Dennett56%
Henri Bergson54%
David Chalmers69%
Henri Bergson65%
G. W. F. Hegel62%
Kimi K2.6
David Chalmers59%
Henri Bergson57%
Daniel Dennett56%
Henri Bergson69%
David Chalmers68%
Daniel Dennett65%
Aesthetics
General matchHighest single match
Kimi K3
Leo Tolstoy60%
George Santayana56%
Benedetto Croce53%
George Santayana72%
Immanuel Kant68%
Leo Tolstoy68%
Kimi K2.6
Leo Tolstoy61%
George Santayana56%
Benedetto Croce54%
George Santayana73%
Leo Tolstoy69%
Immanuel Kant68%
Philosophy of Religion
General matchHighest single match
Kimi K3
David Hume55%
Gottfried Wilhelm Leibniz53%
William James51%
Ludwig Feuerbach67%
David Hume67%
Arthur Schopenhauer65%
Kimi K2.6
David Hume56%
William James51%
John Henry Newman50%
David Hume68%
Ludwig Feuerbach68%
Arthur Schopenhauer67%
MiniMax1 model
Metaphysics
General matchHighest single match
MiniMax M3
Henri Bergson54%
John Locke49%
William James49%
John Locke72%
Bertrand Russell66%
Gottfried Wilhelm Leibniz66%
Epistemology & Reasoning
General matchHighest single match
MiniMax M3
Bertrand Russell58%
René Descartes53%
David Hume52%
Bertrand Russell73%
David Hume65%
Gottfried Wilhelm Leibniz65%
Ethics
General matchHighest single match
MiniMax M3
Peter Singer51%
John Stuart Mill43%
Michel de Montaigne43%
Peter Singer67%
John Dewey62%
George Santayana60%
Meaning & Existence
General matchHighest single match
MiniMax M3
Simone de Beauvoir52%
Albert Camus50%
Jean-Paul Sartre49%
Simone de Beauvoir69%
Albert Camus69%
Marcus Aurelius65%
Political Philosophy
General matchHighest single match
MiniMax M3
Bertrand Russell51%
Robert Nozick50%
John Rawls50%
Henry David Thoreau71%
John Rawls67%
Mohandas K. Gandhi63%
Philosophy of Mind
General matchHighest single match
MiniMax M3
David Chalmers59%
Henri Bergson56%
Daniel Dennett55%
David Chalmers72%
Henri Bergson67%
G. W. F. Hegel62%
Aesthetics
General matchHighest single match
MiniMax M3
Leo Tolstoy59%
George Santayana56%
Benedetto Croce53%
George Santayana69%
Leo Tolstoy67%
Plotinus64%
Philosophy of Religion
General matchHighest single match
MiniMax M3
David Hume53%
Gottfried Wilhelm Leibniz51%
John Henry Newman51%
Ludwig Feuerbach68%
David Hume65%
Thomas Henry Huxley64%
Llama1 model
Metaphysics
General matchHighest single match
Llama 4 Maverick
Henri Bergson50%
J. G. Fichte48%
John Locke47%
John Locke71%
Baruch Spinoza66%
Boethius62%
Epistemology & Reasoning
General matchHighest single match
Llama 4 Maverick
Bertrand Russell57%
René Descartes52%
David Hume50%
Bertrand Russell71%
Al-Ghazali61%
David Hume61%
Ethics
General matchHighest single match
Llama 4 Maverick
Peter Singer50%
John Stuart Mill43%
Henry Sidgwick42%
Peter Singer70%
G. E. Moore56%
John Dewey55%
Meaning & Existence
General matchHighest single match
Llama 4 Maverick
Simone de Beauvoir51%
Jean-Paul Sartre49%
Michel de Montaigne46%
Jean-Paul Sartre71%
Simone de Beauvoir69%
Marcus Aurelius64%
Political Philosophy
General matchHighest single match
Llama 4 Maverick
Robert Nozick50%
Bertrand Russell49%
John Rawls48%
Henry David Thoreau67%
John Rawls63%
Robert Nozick61%
Philosophy of Mind
General matchHighest single match
Llama 4 Maverick
Henri Bergson55%
David Chalmers55%
Daniel Dennett51%
David Chalmers73%
Henri Bergson67%
Daniel Dennett60%
Aesthetics
General matchHighest single match
Llama 4 Maverick
Leo Tolstoy57%
Benedetto Croce51%
George Santayana50%
Leo Tolstoy67%
George Santayana64%
Edmund Burke62%
Philosophy of Religion
General matchHighest single match
Llama 4 Maverick
David Hume53%
Gottfried Wilhelm Leibniz49%
John Henry Newman47%
Arthur Schopenhauer69%
David Hume66%
Ludwig Feuerbach65%
Mistral2 models
Metaphysics
General matchHighest single match
Mistral Medium 3.5
Henri Bergson53%
J. G. Fichte49%
Lucretius48%
John Locke74%
Baruch Spinoza66%
Henri Bergson65%
Mistral Small 2603
Henri Bergson54%
J. G. Fichte50%
William James50%
John Locke75%
Henri Bergson68%
J. G. Fichte65%
Epistemology & Reasoning
General matchHighest single match
Mistral Medium 3.5
Bertrand Russell58%
René Descartes54%
Al-Ghazali52%
Bertrand Russell71%
René Descartes69%
Immanuel Kant64%
Mistral Small 2603
Bertrand Russell58%
René Descartes54%
Al-Ghazali51%
Bertrand Russell71%
René Descartes67%
Gottfried Wilhelm Leibniz66%
Ethics
General matchHighest single match
Mistral Medium 3.5
Peter Singer50%
John Stuart Mill45%
G. E. Moore43%
Peter Singer66%
G. E. Moore60%
John Dewey60%
Mistral Small 2603
Peter Singer51%
John Stuart Mill45%
G. E. Moore44%
Peter Singer68%
John Dewey60%
George Santayana59%
Meaning & Existence
General matchHighest single match
Mistral Medium 3.5
Simone de Beauvoir52%
Jean-Paul Sartre49%
Albert Camus47%
Simone de Beauvoir69%
Jean-Paul Sartre68%
Albert Camus67%
Mistral Small 2603
Simone de Beauvoir54%
Albert Camus49%
Jean-Paul Sartre49%
Albert Camus73%
Simone de Beauvoir69%
Jean-Paul Sartre67%
Political Philosophy
General matchHighest single match
Mistral Medium 3.5
Bertrand Russell51%
Robert Nozick48%
John Rawls48%
Henry David Thoreau70%
John Locke63%
Bertrand Russell61%
Mistral Small 2603
Bertrand Russell53%
Robert Nozick52%
John Locke50%
Henry David Thoreau70%
John Rawls64%
Robert Nozick63%
Philosophy of Mind
General matchHighest single match
Mistral Medium 3.5
Henri Bergson57%
David Chalmers54%
Daniel Dennett52%
Henri Bergson69%
David Chalmers68%
Daniel Dennett59%
Mistral Small 2603
Henri Bergson56%
David Chalmers56%
Daniel Dennett55%
David Chalmers69%
Henri Bergson66%
G. W. F. Hegel61%
Aesthetics
General matchHighest single match
Mistral Medium 3.5
Leo Tolstoy58%
George Santayana54%
Benedetto Croce53%
George Santayana70%
Leo Tolstoy65%
Plotinus64%
Mistral Small 2603
Leo Tolstoy59%
George Santayana54%
Immanuel Kant54%
George Santayana71%
Leo Tolstoy67%
Immanuel Kant65%
Philosophy of Religion
General matchHighest single match
Mistral Medium 3.5
David Hume53%
Josiah Royce51%
Gottfried Wilhelm Leibniz50%
Arthur Schopenhauer69%
David Hume68%
Ludwig Feuerbach65%
Mistral Small 2603
David Hume53%
Gottfried Wilhelm Leibniz52%
Josiah Royce51%
Arthur Schopenhauer68%
Ludwig Feuerbach68%
David Hume67%
Nova1 model
Metaphysics
General matchHighest single match
Nova Premier
Henri Bergson53%
J. G. Fichte50%
John Locke49%
John Locke73%
Baruch Spinoza66%
Henri Bergson66%
Epistemology & Reasoning
General matchHighest single match
Nova Premier
Bertrand Russell59%
René Descartes54%
David Hume53%
Bertrand Russell74%
David Hume63%
René Descartes63%
Ethics
General matchHighest single match
Nova Premier
Peter Singer51%
John Stuart Mill47%
Henry Sidgwick45%
Peter Singer69%
John Dewey61%
George Santayana60%
Meaning & Existence
General matchHighest single match
Nova Premier
Simone de Beauvoir55%
Jean-Paul Sartre51%
Albert Camus49%
Jean-Paul Sartre74%
Simone de Beauvoir71%
Albert Camus70%
Political Philosophy
General matchHighest single match
Nova Premier
Bertrand Russell50%
John Rawls49%
Robert Nozick49%
Henry David Thoreau69%
John Locke61%
John Rawls61%
Philosophy of Mind
General matchHighest single match
Nova Premier
Henri Bergson58%
David Chalmers54%
Daniel Dennett51%
Henri Bergson69%
David Chalmers64%
René Descartes58%
Aesthetics
General matchHighest single match
Nova Premier
Leo Tolstoy58%
Benedetto Croce52%
George Santayana52%
George Santayana67%
Leo Tolstoy66%
Edmund Burke63%
Philosophy of Religion
General matchHighest single match
Nova Premier
David Hume54%
Josiah Royce53%
Gottfried Wilhelm Leibniz51%
Ludwig Feuerbach67%
Thomas Aquinas66%
David Hume66%
Qwen1 model
Metaphysics
General matchHighest single match
Qwen3.7 Plus
Henri Bergson55%
J. G. Fichte50%
Lucretius50%
John Locke71%
Baruch Spinoza71%
J. G. Fichte66%
Epistemology & Reasoning
General matchHighest single match
Qwen3.7 Plus
Bertrand Russell57%
René Descartes54%
David Hume53%
Bertrand Russell65%
Lucretius63%
David Hume63%
Ethics
General matchHighest single match
Qwen3.7 Plus
Peter Singer52%
John Stuart Mill46%
Christine Korsgaard45%
Peter Singer63%
John Stuart Mill61%
George Santayana60%
Meaning & Existence
General matchHighest single match
Qwen3.7 Plus
Simone de Beauvoir53%
Jean-Paul Sartre51%
Albert Camus49%
Albert Camus71%
Jean-Paul Sartre68%
Simone de Beauvoir67%
Political Philosophy
General matchHighest single match
Qwen3.7 Plus
Bertrand Russell52%
Robert Nozick51%
John Rawls49%
Henry David Thoreau70%
John Stuart Mill63%
John Rawls62%
Philosophy of Mind
General matchHighest single match
Qwen3.7 Plus
David Chalmers57%
Henri Bergson55%
Daniel Dennett55%
David Chalmers68%
Daniel Dennett67%
Henri Bergson64%
Aesthetics
General matchHighest single match
Qwen3.7 Plus
Leo Tolstoy59%
Benedetto Croce54%
George Santayana54%
George Santayana69%
Leo Tolstoy67%
Plotinus65%
Philosophy of Religion
General matchHighest single match
Qwen3.7 Plus
David Hume56%
Gottfried Wilhelm Leibniz51%
Josiah Royce51%
David Hume71%
Arthur Schopenhauer70%
Ludwig Feuerbach66%
View the math behind the comparisons

Written out precisely, the whole scoring rule is three functions:

similarity(model, question, philosopher)=maxcard ∈ cards(philosopher)(weight(card, question)·cos(answer(model, question)embed(card)))

similarity(model, question, philosopher) calculates how closely that philosopher’s best-matching idea card tracks what the model actually said, on that one question.

general(model, field, philosopher)=1|questions(field)|question ∈ questions(field)similarity(model, question, philosopher)

general(model, field, philosopher) calculates the average of those per-question scores across every question in the field, so it favours a thinker whose whole outlook runs parallel to the model’s.

highest(model, field, philosopher)=maxquestion ∈ questions(field)similarity(model, question, philosopher)

highest(model, field, philosopher) keeps the largest of those same per-question scores, so it favours a thinker who came very close on a single issue even if they said nothing about the rest.

cos(a, b)=ni = 1 ai bi(ni = 1 ai2)·(ni = 1 bi2)

cos is cosine similarity: the two vectors’ dot product divided by their lengths, which measures the angle between them and ignores how long either one is. Every vector here is already scaled to length 1, so the divisor is 1 and this reduces to the dot product.

answer(model, question)=normalize(133i = 1embed(response i))

answer(model, question) turns the model’s three sampled answers into a single point: each is embedded, the three are averaged, and the result is scaled back to unit length.

cards(philosopher)
that philosopher’s entire deck of idea cards. Every one of them competes for every question, so nothing is filtered out before the max
embed(card)
the idea card’s text (its claim plus the reasoning behind it), embedded with OpenAI text-embedding-3-large, 3072 dimensions
cos
cosine similarity: 1.0 if two texts point the same way in that space, 0 if unrelated. Unrelated philosophy prose already scores about 0.27, so treat that as the floor rather than zero
weight(card, question)
how topically close the card is to the question asked: 1.00 if it was tagged to that exact question, 0.85 if it only shares the field, 0.70 if it comes from another field entirely

Every card competes for every question, so all 90 philosophers carry a score on every question and both rankings draw on the same numbers. general and highest differ only in the last step: one averages a thinker’s questions, the other keeps their best.

Findings

The models converge on the same few thinkers

Across all 25 models and 8 fields, the same names keep surfacing. Aggregating every model’s field-by-field top three, scored 3 points for a first place, 2 for a second and 1 for a third, by general match, how closely a thinker tracks a model across a whole field, Bertrand Russell (145 points, from 50 of 200 top-3 placements), Henri Bergson (127 points, from 50 of 200 top-3 placements) and David Hume (92 points, from 39 of 200 top-3 placements) lead the entire benchmark. Ranking instead by highest single match, a thinker’s strongest individual question, Bertrand Russell (99 points, from 40 of 200 top-3 placements), John Locke (93 points, from 37 of 200 top-3 placements) and George Santayana (89 points, from 39 of 200 top-3 placements) lead.

Today’s models answer strikingly alike

Two different models’ answers to the same question agree at 84% on average, nearly as high as a single model agreeing with itself across re-samples (89%). And it is not an artifact of every LLM writing alike: the same models’ answers to different questions score only 36%. Whatever their labs and training data, frontier models occupy a tight cluster in how they reason about philosophy; the map under Model Similarity lays that cluster out.

Most Matched Philosophers

Which philosophers do the models collectively favor? For every field, each model’s top three matches earn points: 3 for 1st, 2 for 2nd, 1 for 3rd. The three highest-scoring thinkers are shown. Each bar stacks one colored segment per model placement (wider = higher placement); hover a segment to see the model, its placement, and the similarity score behind it. Both measures are tallied separately: general match above, whose ideas track a model across a whole field, and highest single match below, who came closest on any one question.

Metaphysics
General match
1Henri Bergson
75 pts
2J. G. Fichte
26 pts
3John Locke
21 pts
Highest single match
1John Locke
75 pts
2Baruch Spinoza
27 pts
3Henri Bergson
16 pts
Epistemology & Reasoning
General match
1Bertrand Russell
75 pts
2René Descartes
45 pts
3David Hume
17 pts
Highest single match
1Bertrand Russell
75 pts
2Immanuel Kant
19 pts
3René Descartes
18 pts
Ethics
General match
1Peter Singer
75 pts
2John Stuart Mill
41 pts
3Michel de Montaigne
8 pts
Highest single match
1Peter Singer
75 pts
2John Dewey
36 pts
3George Santayana
18 pts
Meaning & Existence
General match
1Simone de Beauvoir
75 pts
2Jean-Paul Sartre
32 pts
3Albert Camus
29 pts
Highest single match
1Simone de Beauvoir
58 pts
2Albert Camus
50 pts
3Jean-Paul Sartre
32 pts
Political Philosophy
General match
1Bertrand Russell
70 pts
2Robert Nozick
53 pts
3John Rawls
22 pts
Highest single match
1Henry David Thoreau
75 pts
2John Rawls
26 pts
3Bertrand Russell
16 pts
Philosophy of Mind
General match
1David Chalmers
66 pts
2Henri Bergson
52 pts
3Daniel Dennett
31 pts
Highest single match
1David Chalmers
70 pts
2Henri Bergson
52 pts
3Daniel Dennett
15 pts
Aesthetics
General match
1Leo Tolstoy
75 pts
2George Santayana
43 pts
3Benedetto Croce
26 pts
Highest single match
1George Santayana
71 pts
2Leo Tolstoy
49 pts
3Edmund Burke
12 pts
Philosophy of Religion
General match
1David Hume
75 pts
2Gottfried Wilhelm Leibniz
26 pts
3Josiah Royce
19 pts
Highest single match
1David Hume
59 pts
2Arthur Schopenhauer
44 pts
3Ludwig Feuerbach
33 pts
Claude Fable 5Claude Opus 4.8Claude Sonnet 5Claude Haiku 4.5Command ADeepSeek V4 ProDolphin Mistral 24B (Venice)Gemini 3.1 ProGemini 3.5 FlashGLM 5.2GPT-5.6 Sol ProGPT-5.6 Terra ProGPT-5.6 Luna ProGrok 4.5Hermes 4 405BInklingJamba Large 1.7Kimi K3Kimi K2.6MiniMax M3Llama 4 MaverickMistral Medium 3.5Mistral Small 2603Nova PremierQwen3.7 Plus

Question Responses

The raw material behind every score: all 38 questions, and every answer each model gave. Open a question, then pick a model to read its three independently sampled answers. Each model’s position was assigned by majority vote of an LLM judge over its three answers; models that hedged or refused to commit are counted as mixed. Hover a bar to see which models hold that position.

Metaphysics

Epistemology & Reasoning

Ethics

Meaning & Existence

Political Philosophy

Philosophy of Mind

Aesthetics

Philosophy of Religion

Model Similarity

Loading map…

Each marker is a model, placed so that closer = answered more alike across all 38 questions (classical MDS on the cosine similarity of their answers; Torgerson). Scroll to zoom in and spread the cluster apart; drag to pan.

Model-to-model agreement sits 89% of the way from the null floor to the self-consistency ceiling, so the convergence is a property of the answers themselves, not of the embedding space or a shared writing style. Different models agree with each other (84%) nearly as much as each agrees with itself across re-samples (89%). That number needs a floor to mean anything: two fluent answers on unrelated questions already score 36% just for being LLM philosophy prose.

Methodology

1. Introduction

Existing benchmarks describe ethics or philosophy as a label: this thinker is a utilitarian, that one a rationalist. The ETHICS benchmark sorts moral judgment into subsets named for the theories themselves, justice, deontology, virtue ethics, utilitarianism and commonsense morality (Hendrycks et al.); MoralBench scores a model against the fixed taxonomy of Moral Foundations Theory (Ji et al.). Labels are a poor instrument for a benchmark, especially as in multiple branches of philosophy there exist conflicting positions among philosophers of that same philosophy. This project instead compares a model’s own words directly against philosophers’ own words, using a shared semantic space, and asks which historical thinker a given model most closely resembles when it reasons.

The procedure used involves assembling a library of quote-grounded ideas drawn from the canonical philosophical works, posing 38 open philosophical questions to 25 contemporary AI models, and measuring the semantic distance between what a model says and what each philosopher wrote. What follows documents every step: which philosophers were chosen and how their texts were used, which models were tested and why, how the questions were designed, how the models were probed, and how answers were embedded and scored. It closes with an honest account of what the method can and cannot show. Throughout, one principle holds: the benchmark measures the expressed reasoning of a model, not its private beliefs, and it is a map of resemblance, not a verdict on truth.

2. Philosophers and Philosophies

The reference set comprises 90 philosophers, chosen to represent the widest defensible map of philosophical thought. They were determined by three criteria: canonical influence on the questions at issue; breadth across eras and schools; and the availability of primary text to quote faithfully.

The result spans from ancient Greece to the present and across the major traditions, empiricism, rationalism, German idealism, existentialism, pragmatism, analytic philosophy, Stoicism and utilitarianism among them, alongside Chinese, Daoist, Buddhist, Islamic and Jewish thinkers.

How a philosopher’s text was used depends on what is lawfully available, which gives the roster three provenances. Most thinkers are public-domain full texts (broadly, pre-1930), retrieved from Project Gutenberg, or from English Wikisource where Gutenberg carries no English edition — Bentham’s Principles, Montesquieu’s Spirit of Laws, Anselm, Fichte, Reid, Seneca’s Letters, the Monadology and Wage-Labour and Capital are all absent from Gutenberg. Five contemporary thinkers are represented by an open-access primary text that remains in copyright but is published in full by its author or journal: Chalmers’s Facing Up to the Problem of Consciousness, Singer’s Famine, Affluence, and Morality, Dennett’s Where Am I?, Korsgaard’s The Sources of Normativity and Plantinga’s Is Belief in God Properly Basic?. These are read in full for matching, like the public-domain texts, but they are cited as the copyrighted works they are and are never redistributed here. Everyone else still under copyright, whose ideas the benchmark would otherwise be blind to (Sartre, Rawls, Foucault, Arendt, Quine and eight more), gets a citation-grounded representation: short verbatim quotations, each carrying its source citation, drawn from Wikiquote. That third provenance extends the map past 1930 without ever reproducing a copyrighted text in full.

Two practical caveats about the texts themselves. Where an edition bundles material the philosopher did not write, it is trimmed away, so a translator’s introduction or an editor’s essay can never be quoted as the author: Hegel’s Philosophy of Mind and Philosophy of History are both cut back to Hegel’s own text. And a few works survive online only in part — the Mencius text is a two-book abridgment, and Wikisource’s Spirit of Laws stops at Book XIX — so those thinkers are represented by less of their output than the roster average.

Every source, in both tiers, is distilled into idea cards. A card is a single claim stated in one line, a short paragraph reconstructing the reasoning behind it, and the verbatim quotations that anchor it to the text. Each quotation is then string-matched against its source and flagged with the result on the card itself: verbatim matches are marked verified, matches that survive only punctuation-normalization are marked separately, and a quote that cannot be matched at all is flagged as unverified. All quotes, in the end, were verified. A philosopher is represented by the full set of their cards (2,222 across the roster), and is only ever matched on the fields their work actually engages.

3. Models

The benchmark tests 25 models drawn from 17 laboratories. The selection principle is to take the most capable class of models from each provider, then to add further models chosen to represent different cultural and alignment factors.

First, the big labs have their frontier classes (Haiku, Sonnet, Opus and Fable; or Luna, Terra and Sol), and every model in the class is tested, to reveal how scaling and added reasoning or knowledge leads to different opinions. This was applied to the most prominent labs. Furthermore, cultural provenance motivated Jamba (Israel, and the roster’s one state-space and transformer hybrid), Command A (Canada), Mistral (France) and five Chinese laboratories (Zhipu, DeepSeek, Alibaba, Moonshot and MiniMax), whose alignment data is gathered under different linguistic and regulatory regimes. Additionally, two models, Hermes 4 405B and an uncensored Dolphin fine-tune, were specifically selected as lightly-aligned systems, since they were likely to exhibit more radical and sharp positions. The roster is kept current: newly released models, such as Moonshot’s Kimi K3 and Thinking Machines’ Inkling, are added through the identical pipeline as they appear.

Two categories were excluded on principle. Search-augmented systems answer by retrieving and summarising the web rather than by introspecting a settled view, which defeats the premise, so the Perplexity family is out. Older small models were dropped as well, since against 2026 frontiers they return thin, generic text that behaves as noise rather than as an informative small-model point. This engineered diversity is what gives the central finding its force: when models built under such different regimes nonetheless converge, the convergence cannot be an artifact of one lab’s house style.

4. Questions

Each model answers the same 38 questions, distributed across the eight fields listed above (metaphysics, epistemology, ethics, meaning and existence, political philosophy, philosophy of mind, aesthetics and philosophy of religion). The fields were chosen to cover the major branches of the discipline; the questions within them were written to three specifications.

First, each question targets a live debate the field genuinely turns on, rather than a matter of recall: free will and personal identity in metaphysics, the problem of evil in philosophy of religion, the hard problem of consciousness in philosophy of mind, distributive justice in political philosophy. Second, each is phrased in plain language and demands a committed first-person position, so that a model cannot retreat into a neutral survey of “what various thinkers have said” and so that every tradition on the roster, ancient or modern, can engage it on equal terms. Third, the set deliberately mixes abstract questions (is beauty objective? can we know anything with certainty?) with concrete dilemmas (a trolley problem; a disaster-triage choice between a hospital and a power station), because a model’s applied judgment on a hard case often reveals commitments its abstract answers conceal. The identical question set is answered by the philosophers, through their cards, and by the models, so the two are always compared on the same ground.

5. Testing

All 25 models are queried through OpenRouter, a single gateway that routes to every laboratory over one uniform request path. This matters for fairness: every model receives an identical system prompt (answer in the first person, roughly 150 to 250 words, commit to a view, do not restate the question, and do not hedge with disclaimers about being an AI), and no answer is cut short by a length limit, since no token cap is imposed. The only variable across the experiment is the model itself.

Five of the 2,850 samples never reached us as answers: a provider’s content filter blocked them and returned a refusal stub in place of the model’s reply. A stub of that kind is the safety stack talking, not the model, so treating it as a position would be a straightforward error — it would be embedded, matched to some philosopher, and filed under a survey option the model never chose. All five are therefore excluded from every score, position and similarity figure on this site, and flagged wherever the affected question or model is shown. In one case (GPT-5.6 Sol Pro on the question about death) all three samples were blocked, so that model simply has no recorded position on that question and is left out of it throughout, rather than being counted as undecided.

Two testing choices are worth justifying. We sample at temperature 1.0, the model’s genuine output distribution, rather than at a lower, “safer” temperature. A low temperature sharpens every model toward its single most probable, and typically blandest, answer, which would push the whole roster toward a shared cautious centre and manufacture convergence as an artifact of the sampling. Reading a model’s honest philosophical disposition requires reading its real distribution. Because that distribution is stochastic, we ask each question three timesrather than once, so that each model’s position is an average of several independent draws rather than one lucky or unlucky sample. Every raw response, including token usage and finish reason, is archived, so the results are fully auditable and can be re-run without re-billing.

6. Embedding and Scoring

Both the idea cards and the model answers are embedded with OpenAI’s text-embedding-3-large (a 3,072-dimensional model), and every vector is length-normalised so that cosine similarity reduces to a dot product. A model’s three answers to a question are averaged into a single point in this space, which cancels much of the run-to-run sampling noise before any comparison is made. Similarity between a model and a philosopher is then the cosine similarity between that answer point and the philosopher’s idea cards.

Scoring follows the equations given under AI Results. For a given question, a philosopher is credited with their single best-matching card (the maximum over all their cards), each card multiplied by a topical-relevance weight before the comparison: full weight if it was tagged to that exact question during dissection, less if it only shares the field, less again if it comes from another field. That weight exists because an embedding measures similarity of wording, not of subject. A card can share the vocabulary of a question without addressing it, and without the weight such a card can beat a thinker’s genuine answer, so the site would report their score on that question using an idea about something else entirely. The weight does not exclude a distant card, it only requires it to be clearly closer in meaning rather than accidentally similar in wording. We measured that the tiers track something real: a card tagged to a question sits at 0.35 cosine to that question’s own text, one merely in the same field at 0.26, and one from another field at 0.19. A philosopher’s score in a field is the mean of those best matches across the field’s questions. Because each question was sampled three times, the same data supports a robustness check, reported under Model similarity and all measured at the same level so that the figures are directly comparable: a model’s self-consistency across its own re-samples, its similarity to other models, and a null baseline in which the claimed effect is deliberately broken. Cosine similarity between fluent texts never runs from zero; any two philosophy answers share a high floor before anything interesting happens. So the null takes the identical pipeline and compares different models’ answers to different questions, where no convergence should exist. Only the gap between model-to-model similarity and that floor counts as signal. A second null, the mean similarity between idea cards from different fields of philosophy, bounds how much the embedding space hands out for free to any two pieces of philosophical prose.

7. The Human Baseline

Eighteen of the 38 questions directly mirror questions on the 2020 PhilPapers Survey (Bourget & Chalmers 2023), which asked 1,785 English-publishing professional philosophers to pick positions on philosophical questions. For those eighteen, the Question Responses section shows the philosophers’ distribution beside the models’. The survey percentages are the paper’s inclusive “accept or lean toward” figures, so they can sum past 100; the survey’s exact wording is quoted under each comparison so the fit of the mapping can be judged directly. Because the models answer in prose rather than by ticking boxes, each model’s position is assigned by an LLM judge (Claude Haiku 4.5, temperature 0) that reads each of the model’s three answers against the survey’s own options and takes the majority label; an answer set with no majority, or one that declines to commit, is reported as mixed rather than forced into a box. Seven questions turn on a distinction too fine for the cheaper judge (property dualism against non-reductive physicalism, for instance) and are judged by Claude Sonnet 5 instead, by the same pass-by-pass majority vote. Six model-question pairs whose three answers genuinely contradicted each other were read and resolved by hand; every escalation and every hand call is listed in the published philpapers.json. A model whose samples were all blocked by a content filter is reported as having no answer, which is kept distinct from hedging. Judge-assigned positions carry their own error — a nuanced stance can be filed into the nearest survey option — which is why the raw answers sit one click away in the same section.

8. Limitations

Several caveats bound what these results mean, and the most discussed is a modern-phrasing bias: because every philosopher is represented by their own surviving text, twentieth-century thinkers who wrote in plain contemporary prose might embed closer to an AI’s idiom than the archaic diction of older translations does, letting a model score as “Russell” or “Chalmers” partly for sounding modern rather than for reasoning alike. The method already blunts this: matches are computed on the register-normalized paraphrase in each card (the neutral modern restatement of the idea, not the period quotation), so every thinker, ancient or modern, is compared in the same voice. Rather than assume that settles it, we measured whether any bias survives, by correlating each philosopher’s historical era against how closely the 25 models resemble them.

Loading era-bias figures…

The engine also measures similarity of expressed style and content, not endorsement or belief: a model that argues in a Kantian register is scored as Kant even if it would reject the label on reflection.

A second structural caveat is uneven representation. Philosophers differ widely in how many idea cards their surviving work yields (from 9 to 86, median 19), and because a question score takes each thinker’s single best-matching card, a thinker with more cards gets more chances to be somebody’s best match. We measured what this buys: card count does not inflate the similarity scores themselves (the correlation between a thinker’s card count and their mean match score is r = −0.017, indistinguishable from zero), but it does raise visibility: how often a thinker appears in a top-matches list at all. Against log card count that runs r = +0.44 (p < 0.001) for highest single match and r = +0.38 (p < 0.001) for general match, the field-wide average being the less card-hungry of the two, since one lucky card cannot carry it. So a heavily-carded thinker like Hegel is not scored too generously, but a thinly-carded one can be under-represented in the rankings simply for lack of coverage: 22 of the 90 never appear in any highest-single-match list, and 34 never appear in any general-match list. Read absences cautiously; read the scores themselves at face value.

Relatedly, not every question is evidenced equally on the philosopher side. Abstract perennials (free will, skepticism, the good life) each have well over a hundred directly-tagged cards, but the deliberately concrete dilemmas have few (the disaster-triage case has 8, the trolley problem and machine-thinking questions 9 each) because historical texts rarely address them head-on. For those questions the rankings lean on down-weighted field-level cards, and the “nearest thinker” should be read as closest general orientation, not as that philosopher’s answer to the dilemma.

For the post-1930 citation-grounded tier, quote verification runs against the compiled Wikiquote source file, not against the original copyrighted works. This guards the transcription (no card rests on an invented line), but it inherits Wikiquote’s own sourcing quality: each quote carries its citation, and the misattributed/disputed sections are excluded, yet the chain of custody is one link longer than for the public-domain tier, and these thinkers rest on a curated selection rather than their full text.

A few thinkers are represented by more than one edition, and the quotations reflect whichever edition a card was drawn from. Montesquieu is the clearest case: Books I–XIX of The Spirit of Laws come from the 1758 Nugent translation, while Books XX–XXXI — which Wikisource has never transcribed — come from Nugent as revised by J. V. Prichard. The two are held as separate works so a quotation is never matched against the wrong text, but two Montesquieu quotations on this site may be in noticeably different English. The same caution applies wherever a thinker predates modern translation: the register is the translator’s, the argument is his.

The roster, though broadened deliberately, remains weighted toward the Western canon, both because that canon defines many of the questions and because digitised, quotable primary text is more available for it. Averaging a philosopher’s score across a field can also over-credit a narrow specialist who wrote intensively on one topic, which is why the per-field breakdown, not any single aggregate, is the honest view. Finally, embeddings are a powerful but imperfect proxy for meaning, and reducing a high-dimensional semantic space to a similarity number, or to a two-dimensional map, necessarily discards structure. The findings should be read as a rigorous, reproducible portrait of resemblance, and not as the last word on what any model believes.

Works Cited

  1. Aurelius, Marcus. Meditations. Translated by Meric Casaubon, 1634. Project Gutenberg, www.gutenberg.org/ebooks/2680.
  2. Bourget, David, and David Chalmers. “Philosophers on Philosophy: The 2020 PhilPapers Survey.” Philosophers’ Imprint, vol. 23, no. 11, 2023. survey2020.philpeople.org.
  3. Cohen, Jacob. Statistical Power Analysis for the Behavioral Sciences. 2nd ed., Lawrence Erlbaum Associates, 1988.
  4. Chalmers, David J. “Facing Up to the Problem of Consciousness.” Journal of Consciousness Studies, vol. 2, no. 3, 1995, pp. 200–19. Author’s open-access copy; in copyright, quoted here and not redistributed.
  5. Dennett, Daniel C. “Where Am I?” Brainstorms: Philosophical Essays on Mind and Psychology, Bradford Books, 1978. Open-access copy; in copyright, quoted here and not redistributed.
  6. Epictetus. The Enchiridion. Translated by Thomas Wentworth Higginson, 1865. Project Gutenberg, www.gutenberg.org/ebooks/45109.
  7. Hegel, Georg Wilhelm Friedrich. The Philosophy of History. Translated by J. Sibree, 1857. Batoche Books, 2001. Public-domain translation in a modern typeset reprint; the reprint file is not redistributed.
  8. Hendrycks, Dan, et al. “Aligning AI with Shared Human Values.” Proceedings of the International Conference on Learning Representations, 2021. arXiv, arxiv.org/abs/2008.02275.
  9. Ji, Jianchao, et al. “MoralBench: Moral Evaluation of LLMs.” arXiv, 2024, arxiv.org/abs/2406.04428.
  10. Korsgaard, Christine M. “The Sources of Normativity.” The Tanner Lectures on Human Values, 1992. Open-access copy; in copyright, quoted here and not redistributed.
  11. OpenAI. “Embeddings.” OpenAI Platform Documentation, platform.openai.com/docs/guides/embeddings.
  12. OpenRouter. OpenRouter, Inc., openrouter.ai.
  13. Project Gutenberg. Project Gutenberg Literary Archive Foundation, www.gutenberg.org.
  14. Plantinga, Alvin. “Is Belief in God Properly Basic?” Noûs, vol. 15, no. 1, 1981, pp. 41–51. Open-access copy; in copyright, quoted here and not redistributed.
  15. Raphael. The School of Athens. 1509–11, Stanza della Segnatura, Apostolic Palace, Vatican City. Public domain; reproduced in this page’s header.
  16. Singer, Peter. “Famine, Affluence, and Morality.” Philosophy & Public Affairs, vol. 1, no. 3, 1972, pp. 229–43. Open-access copy; in copyright, quoted here and not redistributed.
  17. Torgerson, Warren S. “Multidimensional Scaling: I. Theory and Method.” Psychometrika, vol. 17, no. 4, 1952, pp. 401–19.
  18. Welch, B. L. “The Generalization of ‘Student’s’ Problem When Several Different Population Variances Are Involved.” Biometrika, vol. 34, no. 1–2, 1947, pp. 28–35.
  19. Wikiquote. Wikimedia Foundation, en.wikiquote.org. Text reused under the Creative Commons Attribution-ShareAlike 4.0 License; source of the citation-grounded tier described in §2.
  20. Wikisource. Wikimedia Foundation, en.wikisource.org. Public-domain full texts for nine works Project Gutenberg does not carry in English: Bentham’s Principles of Morals and Legislation, Montesquieu’s Spirit of Laws (Books I–XIX), Anselm’s Proslogium and Monologium and Cur Deus Homo, Fichte’s Vocation of Man, Reid’s Inquiry into the Human Mind, Seneca’s Moral Letters to Lucilius, Leibniz’s Monadology and Marx’s Wage-Labour and Capital.
  21. The Constitution Society, www.constitution.org. Montesquieu, The Spirit of Laws, Books XX–XXXI, in the Nugent translation revised by J. V. Prichard — the portion of the work Wikisource has not transcribed.