Plexara research
We test Plexara in public
The performance claims on this site come from controlled studies, and we publish the whole study: the method, the raw run data, and the code that turns one into the other. If a number looks too good, don’t take our word for it. Download the runs and check. When the data kills one of our own ideas, we publish that too.
What we publish with every study
A benchmark you cannot rerun is an ad. Each of ours ships as a working repository, and these four things come with it:
- The report
- A brand-neutral technical report in the open-source project, with a frozen PDF and a snapshot of its data archived on Zenodo under a citable DOI. Other vendors and researchers are welcome to cite it, argue with it, or build on it.
- The raw runs
- Every graded attempt, full transcript, and run manifest, committed to the repository. If you want to know why one answer was marked wrong at 2:14 on a Tuesday, the transcript is in the tree.
- The code
- The harness, the deterministic fixtures and seed generators, the graders, and the analysis scripts. Every table and figure regenerates offline from the committed data. No API key, no network access.
- The protocol
- Decision rules written down before the data comes in, and headline runs pinned to tagged releases so the exact build behind a number stays buildable. When a pre-registered hypothesis fails, the failure stays on the public record next to the runs that killed it.
Why we benchmark
Plexara runs in production for real clients, answering real questions every day, and that field record tells us the approach works. Field experience does not tell you by how much, on which kinds of question, or where to aim the next round of engineering. For that you need a number you did not choose: a controlled study that holds everything constant except the thing under test.
The studies also decide what we build. When the knowledge-use study falsified our pre-registered hypothesis about staleness metadata, we dropped the features that hypothesis would have justified before writing a line of product code, and the same data pushed capture filing and input validation up the roadmap instead.
The knowledge-pollution study broke a prediction of ours again, from the other direction. We planted a wrong fact through our own review queue expecting the dangerous one to be the fact nobody can check, and the data said the opposite: the claims that spread are the ones the platform could have verified against your data before approving them. That is now where the review tooling is aimed.
The graph-completion study is what publishing a dead idea looks like in practice. Its pre-registered instrument kill fired, retiring the grand claim we hoped to make about references between knowledge pages, and the published report leads with the kill beside what survived: references hold the cost of discovery flat as a knowledge base grows a hundredfold, and they are the only route that works when search is off or the reading budget is small.
The research series
Each study in the series is published brand-neutral in the open platform project, archived with its raw run data under a citable DOI, and reproducible offline from the committed attempts. New studies join the series as they are published.
The accuracy study · 2026-07-18
Does the platform make an agent measurably more accurate on your data?
On questions that turn on a business rule, accuracy rose from 42.7% to 98.7%. A companion cold-start experiment taught a fresh install six facts one at a time and watched each question class unlock at its own lesson, and a lifecycle re-run measured a fact taught by one person being reused correctly by a different teammate 98.9% of the time.
The knowledge-use study · 2026-07-26
Does an agent actually use the knowledge the platform delivers?
Agents rely completely on delivered knowledge they cannot re-derive: conventions, definitions, policies. With the company definition delivered, confident fabrication fell from 75% to zero, and capable models re-verified every claim they could check.
The knowledge-pollution study · 2026-08-07
What does a wrong fact cost once it has cleared review?
A wrong fact never out-argued the correct source sitting beside it. It did something quieter: on a small model it suppressed the one query that would have refuted it, and every run that ran that query anyway answered correctly. The frontier-class models we run in production took the wrong answer zero times in 96 runs.
The graph-completion study · 2026-08-10
What do references between knowledge pages buy an agent that has to be complete?
Asked to write complete operational documents, agents followed references between knowledge pages voluntarily, grounded every governing constraint while reading 0.2% of a 5,000-page corpus, and kept discovery cost flat as the corpus grew a hundredfold; without the references, cost roughly doubled and the only failed episode appeared. With search off, references were the only route that worked: 96% of constraints recovered against zero. One pre-registered construct died by its own kill condition, and the kill is published with the result.
Start from the archives
Each Zenodo record holds the frozen report PDF and a snapshot of the raw run data behind it, hosted independently of us and of this site. The open repository holds the harness and every run since.




