the public benchmark protocol
A public and reproducible way to test any nonprofit research tool, ours included, against the one fact that cannot be argued with: the field an organization filed with the IRS.
what this is
This is a question set, a grading rule, and an answer key that points at the exact filing behind every correct answer. You run it against any tool you use, aysra included, and the IRS filing decides who was right. Where a tool cannot point at the line behind a number, you learn that before you decide to trust it.
why we publish it
Behavior over claims has to apply to us as well. If a tool tells you a foundation's median grant was a certain figure, you should be able to find the line that proves it. We would rather hand you the method to check us than ask you to take our word for it. Publishing the test is how we hold ourselves to the standard to which we hold the filings.
the ground rules
Ground truth is the filing, not a human label. Most benchmarks argue over what the right answer is, because a person wrote the answer key. Here the key is a specific line on a specific form for a specific tax year. Either the number matches the filing or it does not.
The answer key points at a field, not a fixed number. Every question resolves to a form, a line, a tax year, and for computed questions a formula and its inputs. Anyone can pull the filing and reproduce the value. The protocol stays correct as filings are added and corrected, and no one has to trust a number we typed once and never revisited.
Tool-agnostic by design. The protocol grades an answer against a filing. It does not rank one product against another.
An honest refusal beats a confident invention. A tool that says it cannot compute a cross-year figure from a single filing is behaving well. A tool that invents a plausible percentage is not. The grading treats these very differently, and on purpose, because a wrong answer delivered with confidence is the one a reader acts on.
a note on grading, since we refuse to grade nonprofits
aysra does not grade, rank, or score nonprofits. Grading a tool's answer against a filed field is a different act. It is a check of fact, not a judgment of worth, and the user does the checking against a filing. We score answers against facts and we do not score organizations against opinions.
how answers are graded
Grade each answer into one of five states.
- Correct. Matches the filed field, or the value computed from the filed fields by the stated formula, and can produce the filing on request.
- Correct but uncited. Right number, no path to the filing. True today, unverifiable tomorrow.
- Partial. Right field with the wrong year, right computation with the wrong inputs, or a range offered where a single value exists.
- Honest refusal. Declines and says why, for example that a cross-year figure cannot come from one filing. This is not a failure.
- Fabricated. A confident answer that does not match the filing, or a number with no filing behind it, also known as a hallucination.
The order that matters runs from correct, to correct but uncited, to honest refusal, to partial, to fabricated. A tool that refuses the questions it cannot answer honestly ranks above a tool that guesses at them.
the four tiers
Questions are grouped by what it takes to answer them honestly. The tier tells you which tools could even be right, and where a tool that reads one filing at a time has to refuse or invent.
-
Tier one, single-field lookup. One value on one filing. Any competent tool should get these.
Example: total revenue for a year (990 or 990-PF, Part I, Line 12).
-
Tier two, single-filing derived. One filing, one computation. This is where a tool that transcribes rather than computes starts to slip.
Example: the median grant a foundation paid in a year (990-PF, Part XV grant list).
-
Tier three, cross-year, one organization. More than one filing for the same entity. A single-filing view cannot do this honestly.
Example: the share of last year's grantees a foundation funded again this year (990-PF, Part XV grant lists across two filings).
-
Tier four, cross-organization. Data spanning many filers. Out of reach for any tool that does not hold the full corpus.
Example: a foundation's payout rate against a set of similar-sized peers (990-PF, Part XIII, across a matched peer set from the corpus).
The sharp edge sits between tier two and tier three. That is where a question still sounds answerable from one document but is not, and where confident invention begins.
the question set
Real and public filers, well-known enough that the filings are easy to find and specific enough that a wrong answer is obviously wrong. Each entry is an answer pointer. Pull the filing and the value is whatever was filed. The examples below use the MacArthur Foundation and the Museum of Contemporary Art Chicago, which you can swap for any filers you prefer.
If one question should anchor the page, it is the renewal rate. It is the point where every single-filing and retrieval-only approach has to refuse or invent, which makes it the sharpest single demonstration of the difference the protocol measures.
how to run it
- Get the filing. It is free and available directly from the IRS. The filing is the answer key, so you are never taking our word for anything.
- Ask the tool the question as written, so results stay comparable.
- Grade the answer against the field using the five states above, and note whether the tool could produce the filing when asked.
- Run it twice. A different answer to the same question is a failure on its own, no matter which answer happened to be right.
what counts as ground truth
The value on the filed form, for the tax year named, as submitted to the IRS. For a lookup that is a single line or for a computed question it is the value the stated formula produces over the stated fields, reproducible by anyone with the same filings.
Where a filing is silent, the honest answer is absent, not zero. A tool that reports a zero for a field an organization never filed has invented a fact. Absence is a real state, and it is not the same as none.
why one filing is not always enough
A tool whose view is a single filing can look up a line, and can sometimes compute within that one filing. It cannot honestly answer a question that needs two years of the same filer, or many filers at once. A renewal rate, a peer comparison, a multi-year trend all fall outside one document. Faced with those, a well-behaved tool should refuse and the failure this protocol is built to surface is the tool that answers anyway, because a fluent, confident, wrong number is the one a reader acts on.
That is the whole reason the answer key points at a field. If the number is real, the field is there. If the field is not there, the number was not either.
what this does not do
The protocol does not rank tools, publish scores for named products, or judge any organization. It measures one thing, whether an answer matches the filing behind it, and it lets you measure that for yourself.
test any tool, including aysra
Raw filing data is always free. The protocol, the question set, and the answer key are public. Run it against whatever you use, and start with us.