the public benchmark protocol
A public and reproducible way to test any nonprofit research tool, ours included, against the one fact that cannot be argued with: the field an organization filed with the IRS.
what this is
This is a question set, a grading rule, and an answer key that points at the exact filing behind every correct answer. It is not a leaderboard and it names no product. You run it against any tool you use, aysra included, and the IRS filing decides who was right. Where a tool cannot point at the line behind a number, you learn that before you decide to trust it.
why we publish it
Behavior over claims has to apply to us as well. If a tool tells you a foundation's median grant was a certain figure, you should be able to find the line that proves it. We would rather hand you the method to check us than ask you to take our word for it. Publishing the test is how we hold ourselves to the standard to which we hold the filings.
the ground rules
Ground truth is the filing, not a human label. Most benchmarks argue over what the right answer is, because a person wrote the answer key. Here the key is a specific line on a specific form for a specific tax year. Either the number matches the filing or it does not.
The answer key points at a field, not a fixed number. Every question resolves to a form, a line, a tax year, and for computed questions a formula and its inputs. Anyone can pull the filing and reproduce the value. The protocol stays correct as filings are added and corrected, and no one has to trust a number we typed once and never revisited.
Tool-agnostic by design. The protocol grades an answer against a filing. It does not rank one product against another.
An honest refusal beats a confident invention. A tool that says it cannot compute a cross-year figure from a single filing is behaving well. A tool that invents a plausible percentage is not. The grading treats these very differently, and on purpose, because a wrong answer delivered with confidence is the one a reader acts on.
a note on grading, since we refuse to grade nonprofits
aysra does not grade, rank, or score nonprofits. Grading a tool's answer against a filed field is a different act. It is a check of fact, not a judgment of worth, and the user does the checking against a filing. We score answers against facts and we do not score organizations against opinions.
how answers are graded
Grade each answer into one of five states.
- Correct. Matches the filed field, or the value computed from the filed fields by the stated formula, and can produce the filing on request.
- Correct but uncited. Right number, no path to the filing. True today, unverifiable tomorrow.
- Partial. Right field with the wrong year, right computation with the wrong inputs, or a range offered where a single value exists.
- Honest refusal. Declines and says why, for example that a cross-year figure cannot come from one filing. This is not a failure.
- Fabricated. A confident answer that does not match the filing, or a number with no filing behind it, also known as a hallucination.
The order that matters runs from correct, to correct but uncited, to honest refusal, to partial, to fabricated. A tool that refuses the questions it cannot answer honestly ranks above a tool that guesses at them.
the four tiers
Questions are grouped by what it takes to answer them honestly. The tier tells you which tools could even be right, and where a tool that reads one filing at a time has to refuse or invent.
- Tier one, single-field lookup. One value on one filing. Any competent tool should get these. Example: total revenue for a year.
- Tier two, single-filing derived. One filing, one computation. This is where a tool that transcribes rather than computes starts to slip. Example: the median grant a foundation paid in a year.
- Tier three, cross-year, one organization. More than one filing for the same entity. A single-filing view cannot do this honestly. Example: the share of last year's grantees a foundation funded again this year.
- Tier four, cross-organization. Data spanning many filers. Out of reach for any tool that does not hold the full corpus. Example: a foundation's payout rate against a set of similar-sized peers.
The sharp edge sits between tier two and tier three. That is where a question still sounds answerable from one document but is not, and where confident invention begins.
the question set
Real and public filers, well-known enough that the filings are easy to find and specific enough that a wrong answer is obviously wrong. Each entry is an answer pointer. Pull the filing and the value is whatever was filed. The examples below use the MacArthur Foundation and the Museum of Contemporary Art Chicago, which you can swap for any filers you prefer.
Tier one. Total revenue in a given year. Answer key: 990 Part I, Line 12. Watch for an approximate figure with no citation, or a number rounded to the nearest million. Highest-compensated officer and title. Answer key: 990 Part VII, with Schedule J where it applies. Watch for the right person paired with an invented compensation figure.
Tier two. Median grant paid by a foundation in a given year. Answer key: 990-PF Part XIV, the median over every grant amount filed that year. Watch for an average reported as a median (the single most common fabrication in this category) and for a round number with no distribution behind it. Program expense ratio for a year. Answer key: 990 Part IX, program service expense divided by total functional expense. Watch for a vague paragraph with no numbers, or a ratio that does not reconcile to the filed lines.
Tier three. Grantee renewal rate from one year to the next. Answer key: 990-PF Part XIV across two consecutive filing years, recipients funded in both divided by the prior year's recipients. Watch for a confident percentage from a tool that only ever saw one filing. This is the clearest tell in the whole set. Program expense share across five years. Answer key: 990 Part IX across five filings, the share for each year and then the trend. Watch for the inputs shown but the trend never computed, or a fabricated trajectory.
Tier four. Payout rate against similar-sized foundations. Answer key: this filer's 990-PF payout compared with a peer set matched on asset band. Watch for a comparison with no disclosed peer set, or a benchmark figure that traces to no filers at all.
If one question should anchor the page, it is the renewal rate. It is the point where every single-filing and retrieval-only approach has to refuse or invent, which makes it the sharpest single demonstration of the difference the protocol measures.
how to run it
- Get the filing. It is free and available directly from the IRS. The filing is the answer key, so you are never taking our word for anything.
- Ask the tool the question as written, so results stay comparable.
- Grade the answer against the field using the five states above, and note whether the tool could produce the filing when asked.
- Run it twice. A different answer to the same question is a failure on its own, no matter which answer happened to be right.
what counts as ground truth
The value on the filed form, for the tax year named, as submitted to the IRS. For a lookup that is a single line or for a computed question it is the value the stated formula produces over the stated fields, reproducible by anyone with the same filings.
Where a filing is silent, the honest answer is absent, not zero. A tool that reports a zero for a field an organization never filed has invented a fact. Absence is a real state, and it is not the same as none.
why one filing is not always enough
A tool whose view is a single filing can look up a line, and can sometimes compute within that one filing. It cannot honestly answer a question that needs two years of the same filer, or many filers at once. A renewal rate, a peer comparison, a multi-year trend all fall outside one document. Faced with those, a well-behaved tool should refuse and the failure this protocol is built to surface is the tool that answers anyway, because a fluent, confident, wrong number is the one a reader acts on.
That is the whole reason the answer key points at a field. If the number is real, the field is there. If the field is not there, the number was not either.
what this does not do
The protocol does not rank tools, publish scores for named products, or judge any organization. It measures one thing, whether an answer matches the filing behind it, and it lets you measure that for yourself.
test any tool, including aysra
Raw filing data is always free. The protocol, the question set, and the answer key are public. Run it against whatever you use, and start with us.