PhilosopherBench 1.1A benchmark that evaluates how frontier AI models think throughout a vast landscape of philosophical questions and positions. Each model answers 38 questions across eight fields, and its responses are matched, by semantic embedding, to the ideas of 90 influential philosophers.
This benchmark measures how a model reasons in philosophy by comparing its own words directly against philosophers’ own words. The works of 90 influential philosophers are compiled, summarized, and their key ideas extracted into short, quote-grounded fact cards, and every model answer is matched against them.
The benchmark uses this method rather than sorting models into schools or categories because labels such as “existentialism” or “cynicism” span many different interpretations, occasionally disagreeing with each other. “Stoicism” is one label, but Marcus Aurelius holds that whether the universe is governed by gods is an open question and that it changes nothing about how one should live: “if it be so that there be no gods, or that they take no care of the world, why should I desire to live in a world void of gods, and of all divine providence?” (Aurelius, bk. 2); Epictetus, another Stoic, treats the matter as settled and foundational, making right belief about a governing providence the whole of piety: the essence of piety is “to form right opinions concerning them, as existing and as governing the universe justly and well” (Epictetus, ch. 31). One label, two incompatible commitments about what the universe is.
The method, put simply, is as follows:
The philosopher a model “reasons like” is the closest match.
This benchmark evaluates 8 fields of philosophy:
Results are below; the full method and the complete question-and-response data follow them.
The benchmark compares the model’s answer to the highest-matching card of a philosopher, with the cards related to the given subject weighted more so that merely similar-sounding prose does not score too high. Each philosopher carries one score on every question in a field, and all 90 are ranked. Those scores are summarized two ways:
The score is calculated based on the cosine similarity between a model’s answers in that field and the philosopher’s ideas: 100% would mean the two point in exactly the same direction in the embedding space, and 0% would mean they are unrelated. A philosopher at 64% is a stronger match than one at 58%; the scores rank how closely a model’s expressed reasoning tracks each thinker.
Each model’s closest three philosophers, field by field, grouped by family. Expand any family to see its models broken down across all eight fields.
Claude4 models
Command1 model
DeepSeek1 model
Dolphin1 model
Gemini2 models
GLM1 model
GPT3 models
Grok1 model
Hermes1 model
Inkling1 model
Jamba1 model
Kimi2 models
MiniMax1 model
Llama1 model
Mistral2 models
Nova1 model
Qwen1 modelWritten out precisely, the whole scoring rule is three functions:
similarity(model, question, philosopher) calculates how closely that philosopher’s best-matching idea card tracks what the model actually said, on that one question.
general(model, field, philosopher) calculates the average of those per-question scores across every question in the field, so it favours a thinker whose whole outlook runs parallel to the model’s.
highest(model, field, philosopher) keeps the largest of those same per-question scores, so it favours a thinker who came very close on a single issue even if they said nothing about the rest.
cos is cosine similarity: the two vectors’ dot product divided by their lengths, which measures the angle between them and ignores how long either one is. Every vector here is already scaled to length 1, so the divisor is 1 and this reduces to the dot product.
answer(model, question) turns the model’s three sampled answers into a single point: each is embedded, the three are averaged, and the result is scaled back to unit length.
Every card competes for every question, so all 90 philosophers carry a score on every question and both rankings draw on the same numbers. general and highest differ only in the last step: one averages a thinker’s questions, the other keeps their best.
Across all 25 models and 8 fields, the same names surface. Aggregating every model’s field-by-field top three, scored 3 points for a first place, 2 for a second and 1 for a third, by general match, how closely a thinker tracks a model across a whole field, Bertrand Russell (145 points, from 50 of 200 top-3 placements), Henri Bergson (127 points, from 50 of 200 top-3 placements) and David Hume (92 points, from 39 of 200 top-3 placements) lead the entire benchmark. Ranking instead by highest single match, a thinker’s strongest individual question, Bertrand Russell (99 points, from 40 of 200 top-3 placements), John Locke (93 points, from 37 of 200 top-3 placements) and George Santayana (89 points, from 39 of 200 top-3 placements) lead.
Two different models’ answers to the same question resemble each other at 84% embedding similarity on average, nearly as high as a single model’s resemblance to its own re-samples (89%). It is not solely attributable to every AI model writing alike: the same models’ answers to different questions score only 36%. This implies that these models cluster in how they reason about philosophy; the map under Model Similarity lays that cluster out. Embedding similarity measures shared framing and reasoning rather than identical conclusions; the conclusions themselves are compared in the next finding.
Which philosophers do the models collectively favor? For every field, each model’s top three matches earn points: 3 for 1st, 2 for 2nd, 1 for 3rd. The three highest-scoring thinkers are shown. Each bar stacks one colored segment per model placement (wider = higher placement); hover a segment to see the model, its placement, and the similarity score behind it. Both measures are tallied separately: general match above, whose ideas track a model across a whole field, and highest single match below, who came closest on any one question.
The raw material behind every score: all 38 questions, and every answer each model gave. Open a question, then pick a model to read its three independently sampled answers. Each model’s position was assigned by majority vote of an AI judge over its three answers; models that hedged or refused to commit are counted as mixed. Hover a bar to see which models hold that position.
Loading map…
Each marker is a model, placed so that closer means answered more alike across all 38 questions (classical MDS on the cosine similarity of their answers; Torgerson). Scroll to zoom in and spread the cluster apart; drag to pan.
Existing benchmarks describe ethics or philosophy as a label: this thinker is a utilitarian, that one a rationalist. The ETHICS benchmark sorts moral judgment into subsets named for the theories themselves, justice, deontology, virtue ethics, utilitarianism and commonsense morality (Hendrycks et al.); MoralBench scores a model against the fixed taxonomy of Moral Foundations Theory (Ji et al.). Labels have benefits, but when it comes to schools of philosophy they are inadequate for truly positioning thinkers, especially as in multiple branches of philosophy there exist conflicting positions among philosophers of that same philosophy. This project instead compares a model’s own words directly against philosophers’ own words, using a shared semantic space, and asks which historical thinker a given model most closely resembles when it answers philosophical questions.
The procedure used involves assembling a library of quote-grounded ideas drawn from the canonical philosophical works, posing 38 open philosophical questions to 25 contemporary AI models, and measuring the semantic distance between what a model says and what each philosopher wrote. The central finding is that the models converge: two different models’ answers resemble each other almost as closely as one model’s re-samples resemble each other, and on the questions shared with the 2020 PhilPapers Survey the models agree far more than professional philosophers do. Throughout, one principle holds: the benchmark measures the expressed reasoning of a model, not its private beliefs, and it is a map of resemblance, not a verdict on truth.
The reference set comprises 90 philosophers, chosen to represent the widest defensible map of philosophical thought. They were determined by three criteria: canonical influence on the questions at issue; breadth across eras and schools; and the availability of primary text to quote.
The roster spans from ancient Greece to the present and across the major traditions, empiricism, rationalism, German idealism, existentialism, pragmatism, analytic philosophy, Stoicism and utilitarianism among them, alongside Chinese, Daoist, Buddhist, Islamic and Jewish thinkers.
How a philosopher’s text was used depends on what is lawfully available, which gives the roster three provenances.
Most thinkers are public-domain full texts (broadly, pre-1930), retrieved from Project Gutenberg, or from English Wikisource where Gutenberg carries no English edition (Bentham’s Principles, Montesquieu’s Spirit of Laws, Anselm, Fichte, Reid, Seneca’s Letters, the Monadology and Wage-Labour and Capital).
Five contemporary thinkers are represented by an open-access primary text that remains in copyright but is published in full by its author or journal: Chalmers’s Facing Up to the Problem of Consciousness, Singer’s Famine, Affluence, and Morality, Dennett’s Where Am I?, Korsgaard’s The Sources of Normativity and Plantinga’s Is Belief in God Properly Basic?. These are read in full for matching, like the public-domain texts, but they are cited as the copyrighted works they are and are never redistributed here.
Everyone else still under copyright, whose ideas the benchmark would otherwise be blind to (Sartre, Rawls, Foucault, Arendt, Quine and eight more), gets a citation-grounded representation: short verbatim quotations, each carrying its source citation, drawn from Wikiquote. That third provenance extends the map past 1930 without ever reproducing a copyrighted text in full.
Texts with editorial additions are trimmed; for example, Hegel’s Philosophy of Mind and Philosophy of History are both cut back to Hegel’s own text. A few works survive online only in part (the Mencius text is a two-book abridgment, and Wikisource’s Spirit of Laws stops at Book XIX); these are represented as fully as the surviving text allows.
Every source, from all three provenances, is distilled into idea cards. The cards were drafted by an AI model, Claude Sonnet 5, working through each work under fixed dissection instructions, which permit a card only for a position the author endorses, never for a view the author raises in order to refute it. A card is a single claim stated in one line, a short paragraph reconstructing the reasoning behind it, and the verbatim quotations from the original text. Each quotation is then matched against its source to be verified. All quotes, in the end, were verified. A philosopher is represented by the full set of their cards (2,222 across the roster), and every card competes for every question, weighted by how topically close it is to the question asked.
The benchmark tests 25 models drawn from 17 laboratories. The selection principle is to take a wide variety of leading AI models, specifically text-based multimodal models and large language models, then to add further models chosen to represent different cultural and alignment factors.
First, the big labs have their frontier classes (Haiku, Sonnet, Opus and Fable; or Luna, Terra and Sol), and every model in the class is tested to investigate how scaling and added reasoning or knowledge leads to different opinions. This was applied to most labs with model classes. Furthermore, cultural provenance motivated Jamba (Israel, and the roster’s one state-space and transformer hybrid), Command A (Canada), and Mistral (France), whose alignment data is gathered under different linguistic and regulatory regimes. Additionally, two models, Hermes 4 405B and an uncensored Dolphin fine-tune, were specifically selected as lightly-aligned systems, since they were likely to exhibit more radical and sharp positions. The roster is kept current: newly released models, such as Moonshot’s Kimi K3 and Thinking Machines’ Inkling, are added through the identical pipeline as they appear. (Roster last updated July 18, 2026.)
Each model answers the same 38 questions, distributed across the eight fields: metaphysics, epistemology, ethics, meaning and existence, political philosophy, philosophy of mind, aesthetics and philosophy of religion. The fields were chosen to cover the major branches of the discipline; the questions within them were written to three specifications.
First, each question targets a live debate in the field: free will and personal identity in metaphysics, the problem of evil in philosophy of religion, the hard problem of consciousness in philosophy of mind, distributive justice in political philosophy. Second, each is phrased in plain language and demands a committed first-person position to probe the model to answer itself rather than retreating to a neutral survey of “what various thinkers have said”. Third, the set deliberately mixes abstract questions (is beauty objective? can we know anything with certainty?) with concrete dilemmas (a trolley problem; a disaster-triage choice between a hospital and a power station), for both diversity and because a model’s applied judgment on a hard case often reveals commitments its abstract answers conceal.
All 25 models are queried through OpenRouter, a model router that routes to every laboratory over one uniform request path. This ensures fairness: every model receives an identical system prompt (answer in the first person, roughly 150 to 250 words, commit to a view, do not restate the question, and do not hedge with disclaimers about being an AI). There is no token limit, so a response can run to any length.
Models are sampled at temperature 1.0 to capture the model’s genuine output distribution, rather than at a lower, “safer” temperature. A low temperature would lean toward the most probable, and typically blandest, answer, which would push the whole roster toward a shared cautious centre and manufacture convergence as an artifact of the temperature.
Every question is sampled three times rather than once, so that each model’s position is an average of several independent draws rather than one outlier sample. Every raw response, including token usage and finish reason, is archived, so the results are fully auditable and can be re-run.
Five of the 2,850 samples never reached us as answers: a provider’s content filter blocked them and returned a refusal stub in place of the model’s reply. Such a stub reflects the provider’s safety system rather than the model, so these samples were discarded. All five are thus also excluded from every score, position and similarity figure on this site, and flagged wherever the affected question or model is shown. In one case (GPT-5.6 Sol Pro on the question about death) all three samples were blocked, so that model simply has no recorded position on that question and is left out of it throughout, rather than being counted as undecided.
Both the idea cards and the model answers are embedded with OpenAI’s text-embedding-3-large (a 3,072-dimensional model), and every vector is length-normalized so that cosine similarity reduces to a dot product. A model’s three answers to a question are averaged into a single point in this space to reduce run-to-run variance. Similarity between a model and a philosopher is then the cosine similarity between that answer point and the philosopher’s idea cards.
Scoring follows the equations given under AI Results. For a given question, a philosopher is credited with their single best-matching card (the maximum over all their cards), each card multiplied by a topical-relevance weight before the comparison: full weight if it was tagged to that exact question during dissection, less if it only shares the field, less again if it comes from another field. That weight exists because an embedding measures similarity of wording, not of subject. A card can share the vocabulary of a question without addressing it, and without the weight such a card can beat a thinker’s genuine answer, so the benchmark would report their score on that question using an idea about something else entirely. Concretely, a same-field card is discounted by 15% and a card from another field by 30% (weights of 1.00, 0.85 and 0.70). A card tagged to a question sits at 0.35 cosine to that question’s own text, while one merely in the same field sits at 0.26, and one from another field at 0.19. A philosopher’s score in a field is the mean of those best matches across the field’s questions.
Each model’s three responses are used to measure a model’s self-consistency across its own re-samples, its similarity to other models, and a null baseline. Cosine similarity between fluent texts never runs from zero; any two philosophy answers share a high floor before anything interesting happens. So the null takes the identical pipeline and compares different models’ answers to different questions, to reveal the floor where no convergence should exist. Only the gap between model-to-model similarity and that floor counts as signal. A second null, the mean similarity between idea cards from different fields of philosophy, bounds how much the embedding space hands out for free to any two pieces of philosophical prose.
Loading embedder comparison…
Eighteen of the 38 questions directly mirror questions on the 2020 PhilPapers Survey (Bourget & Chalmers 2023), which asked 1,785 English-publishing professional philosophers to pick positions on philosophical questions. For those eighteen, the Question Responses section shows the philosophers’ distribution beside the models’. The survey percentages are the paper’s inclusive “accept or lean toward” figures, so they can sum past 100; the survey’s exact wording is quoted under each comparison so the fit of the mapping can be judged directly. Because the models answer in prose rather than by ticking boxes, each model’s position is assigned by an AI judge (Claude Haiku 4.5, temperature 0) that reads each of the model’s three answers against the survey’s own options and takes the majority label; an answer set with no majority, or one that declines to commit, is reported as mixed. Seven questions were too fine for the cheaper judge (property dualism against non-reductive physicalism, for instance) and were judged instead by Claude Sonnet 5, by the same pass-by-pass majority vote. Six model-question pairs whose three answers genuinely contradicted each other were read and resolved by hand; every escalation and every hand call is listed. Judge-assigned positions may carry their own error (a nuanced stance can be filed into the nearest survey option), so the raw answers are available for verification.
No benchmark is perfect in estimating an AI’s position or beliefs. However, they do serve as a useful proxy for determining them. This benchmark is by no means perfect for measuring philosopher positions, but it measures something concrete and checkable: how closely a model’s expressed reasoning resembles each philosopher’s own words, with every number reproducible from the published data.
Several other caveats bound what these results mean. First is the modern-phrasing bias: because every philosopher is represented by their own surviving text, twentieth-century thinkers who wrote in plain contemporary prose might embed closer to an AI’s idiom than the archaic diction of older translations does, letting a model score as “Russell” or “Chalmers” partly for sounding modern rather than for reasoning alike. The method already blunts this: matches are computed on the register-normalized paraphrase in each card (the neutral modern restatement of the idea, not the period quotation), so every thinker, ancient or modern, is compared in the same voice. Rather than assume that settles it, we measured whether any bias survives, by correlating each philosopher’s historical era against how closely the 25 models resemble them.
Loading era-bias figures…
The engine also measures similarity of expressed style and content, not endorsement or belief: a model that argues in a Kantian register is scored as Kant even if it would reject the label on reflection.
A second structural caveat is uneven representation. Philosophers differ widely in how many idea cards their surviving work yields (from 9 to 86, median 19), and because a question score takes each thinker’s single best-matching card, a thinker with more cards gets more chances to be somebody’s best match. We measured what this buys: card count raises both the similarity scores themselves (the correlation between a thinker’s card count and their mean match score is r = +0.61, p < 0.001) and visibility: how often a thinker appears in a top-matches list at all. Against log card count that runs r = +0.44 (p < 0.001) for highest single match and r = +0.38 (p < 0.001) for general match, the field-wide average being the less card-hungry of the two, since one lucky card cannot carry it. Part of this reflects genuine coverage, since a thinker who wrote more on these questions has more relevant ideas to match, but part is mechanical, since the best of many cards is higher on average than the best of a few. So a heavily-carded thinker like Hegel may be scored somewhat generously, and a thinly-carded one can be under-represented in the rankings simply for lack of coverage: 22 of the 90 never appear in any highest-single-match list, and 34 never appear in any general-match list. Read both the high placements of heavily-carded thinkers and the absences of thinly-carded ones with this in mind.
Relatedly, not every question is evidenced equally on the philosopher side. Abstract perennials (free will, skepticism, the good life) each have well over a hundred directly-tagged cards, but the deliberately concrete dilemmas have few (the disaster-triage case has 8, the trolley problem and machine-thinking questions 9 each) because historical texts rarely address them head-on. For those questions the rankings lean on down-weighted field-level cards, and the “nearest thinker” should be read as closest general orientation, not as that philosopher’s answer to the dilemma.
For the post-1930 citation-grounded tier, quote verification runs against the compiled Wikiquote source file, not against the original copyrighted works. This guards the transcription (no card rests on an invented line), but it inherits Wikiquote’s own sourcing quality: each quote carries its citation, and the misattributed/disputed sections are excluded, yet the chain of custody is one link longer than for the public-domain tier, and these thinkers rest on a curated selection rather than their full text.
A few thinkers are represented by more than one edition, and the quotations reflect whichever edition a card was drawn from. Montesquieu is the clearest case: Books I–XIX of The Spirit of Laws come from the 1758 Nugent translation, while Books XX–XXXI — which Wikisource has never transcribed — come from Nugent as revised by J. V. Prichard. The two are held as separate works so a quotation is never matched against the wrong text, but two Montesquieu quotations on this site may be in noticeably different English. The same caution applies wherever a thinker predates modern translation: the register is the translator’s, the argument is his.
The roster, though broadened deliberately, remains weighted toward the Western canon, both because that canon defines many of the questions and because digitised, quotable primary text is more available for it. Averaging a philosopher’s score across a field can also over-credit a narrow specialist who wrote intensively on one topic, which is why the per-field breakdown, not any single aggregate, is the honest view. Finally, embeddings are a powerful but imperfect proxy for meaning, and reducing a high-dimensional semantic space to a similarity number, or to a two-dimensional map, necessarily discards structure. The findings should be read as a rigorous, reproducible portrait of resemblance, and not as the last word on what any model believes.