Which AI answers Vedic astrology questions correctly?

A first public run: 9 models, 16 questions, 140 graded answers. Vedic and KP chart questions only. USD 2.43 spent on Vedika API calls, of a USD 50 cap.

Every model gets the same questions, the same birth data and the same answer format. A fixed rule grader checks each answer against reference chart values computed with Swiss Ephemeris. No AI grades another AI. Vedika tiers answer through the Vedika API, which calculates the chart. The other models answer from the model alone, with no tools. With 16 questions each, gaps of a few points are noise, so read the ranges and not the order.

Scroll ↓

What this run measures

16 Vedic and KP questions in English, Hindi, Bengali, Marathi, Telugu and Tamil (KP in English and Hindi only). Each model answered each question once. Reference values come from a single ephemeris calculation.

What it does not

Western, numerology, Vastu, tarot, intent, emotion, guidance, care and honesty tracks are not run, and neither are most Indian languages. No model has been run both with and without the Vedika API, so this page makes no claim about lift. 33 further Vedika answers, including Western, are left out because no other model answered those questions. Not yet run: Gemini, DeepSeek, Qwen, Llama, Mistral, Sarvam. Partial and not ranked: MiniMax M3 (1M) (2 of 16); Grok 4.6 (1 of 16); Grok 4.7 (3 of 16), which took minutes per answer and mostly timed out.

How outside models ran

Through each provider’s own command-line tool or Claude Code with no tools enabled, on our own subscriptions. Costs shown for them are list-price estimates. Vedika costs are what the API billed.

The leaderboard

Every model answers the same 16 questions. The bar is the pass rate, the dashed line is its 95% range, and the order follows the low end of that range. Click a row for its breakdown; hover a row to find it in the charts below.

#Model

Ranked by the low end of each model's 95% range. With this few questions the ranges overlap heavily, so the order is not a verdict. Cost and tokens are averages per answer. For Vedika tiers cost is the billed amount. For other models a leading ≈ marks an estimate from token counts at list price, which overstates OpenAI rows by roughly half because cached input was priced as fresh; Claude rows have no token data. “Per correct” is total cost divided by correct answers. Latency is the median wall-clock time in the 27 September 2026 runs, through each provider’s command-line tool; it was not recorded for the Vedika API rows and the Claude rows.

View the leaderboard as a data table
Pass rate on the 16 shared questions, with a 95% range (Wilson), average cost and tokens per answer.
ModelProviderRightPass rate95% rangeCost per answerTokens per answerHow measured
GPT-5.6 SolOpenAI13 of 1681.3%57 to 93%~$0.14121,777Subscription CLI, no tools
Claude Sonnet 5Anthropic13 of 1681.3%57 to 93%n/an/aClaude Code subagent, no tools
Vedika SwiftVedika12 of 1675%51 to 90%$0.0289,649Vedika API, billed
GPT-5.6 TerraOpenAI12 of 1675%51 to 90%~$0.07121,210Subscription CLI, no tools
Claude Opus 5.5Anthropic12 of 1675%51 to 90%n/an/aClaude Code subagent, no tools
Vedika StandardVedika11 of 1668.8%44 to 86%$0.02813,667Vedika API, billed
Vedika EcoVedika8 of 1266.7%39 to 86%$0.03312,971Vedika API, billed
Vedika Standard 2.5Vedika10 of 1662.5%39 to 82%$0.0178,311Vedika API, billed
Claude Haiku 4.5Anthropic9 of 1656.3%33 to 77%n/an/aClaude Code subagent, no tools

Skill by skill

Vedic chart facts and KP sub-lord reasoning are the two tracks run so far. A grey cell has fewer than five questions behind it and shows right out of asked, because a percentage on two or three questions says little. Written but not yet run: Western, Numerology, Vastu, Tarot, Intent, Emotion, Guidance, Care and safety, Honesty.

Share of questions answered well, by skill

75% and up50 to 75%25 to 50%under 25%

Grey cells have fewer than five questions and show right out of asked. Blue is 75% and up, lavender 50 to 75%, amber 25 to 50%.

Language by language

India asks in many languages, and often in two at once. Each cell is the share of questions answered well when the question was asked in that language, in its own script. Only these six languages have been run, with two to four questions each, so every cell is a count and not a rate. Not yet run: Gujarati, Urdu, Kannada, Odia, Malayalam, Punjabi, Assamese, Hinglish, Tanglish.

Share of questions answered well, by the language of the question

75% and up50 to 75%25 to 50%under 25%

Cells show right out of asked. An answer counts only if it is right and written in the language and script the person used, with the correct local names for signs, stars and planets.

Score against cost

Higher is more accurate; further left is cheaper. The dashed line joins the best-value models: nothing beats them on both score and cost.

Pass rate against average cost per answer

VedikaOther providersBest-value line

Horizontal axis is logarithmic, so a step to the right is a multiple of the cost.

Where models fail

One square per question and model. Scroll to see which questions every model gets right, which ones split them, and which ones nobody gets.

What a correct answer costs

A cheap model that is often wrong can cost more per correct answer than a dearer one that is usually right. Models with no token data are left out.

Cost per correct answer, in US cents

VedikaOther providers

Total spend on a model divided by the number of answers it got right. Vedika bars are billed amounts; OpenAI bars are list-price estimates that run high.

Spend, batch by batch

Every Vedika API call is paid and ledgered from real receipts. No batch may pass USD 10, and the whole bench stops at USD 50.

Cumulative spend, US dollars

Measured from the run ledger. The other models ran through subscriptions we already hold, so they add no per-answer charge; their cost column is an estimate.

How we test

  • Fixed questions, frozen first

    The questions and grading rules were frozen and dated before any paid run. Changes are dated amendments, never silent edits.

  • Reference values from one ephemeris

    Chart facts are computed with Swiss Ephemeris at the settings the Vedika engine uses. A second independent calculation has not been done yet. Early hand-typed references that disagreed were recomputed, and the change is logged.

  • Graded by rules, not by an AI

    A fixed grader checks each answer against the reference. Hand audits of stored answers found and fixed two grader bugs.

  • Same questions for everyone

    Each model gets the same question text, birth data and answer format. Vedika tiers answer through the Vedika API; other models answer raw, with no Vedika tools. Claude models answered through Claude Code subagents with tools disallowed, eight questions per call, which is a looser control than the command-line runs used for the others.

  • Real spend, hard limits

    USD 10 per batch, USD 50 for the whole bench, tracked from receipts. One run was stopped early after a wallet balance dropped for a reason the ledger did not explain; it was traced to a shared wallet and the bench moved to its own.

  • Masked names

    Vedika tiers are shown by their Vedika names only.

Source: Vedika Bench run records and spend ledger, data frozen 27 September 2026 (93 Vedika answered pairs graded in total, 140 shown here). Logos are trademarks of their owners and are used only to identify each provider.
Engine benchmarksPricingDocs