Benchmark question
This benchmark asked whether an MCP-connected agent could move from included full-text studies to source-backed extraction proposals inside Coventra, without bypassing the review queue.
The goal was workflow readiness: could the agent work from the project protocol, focus on configured outcomes, cite supporting PDF evidence, avoid duplicate proposals on repeat runs, and leave the researcher in control of the final record?
This article reports system behavior and workflow readiness. It is not a clinical accuracy claim and does not replace human verification against the source paper.
Project context
The extraction benchmark used the same oncology SRMA benchmark project as the screening run: first-line PD-(L)1 plus chemotherapy in advanced non-squamous NSCLC.
After title/abstract and full-text progression, the project had 11 included trial reports with PDFs ready for extraction review and 17 protocol outcomes configured before extraction began.
Extraction benchmark starting point
| Item | Observed state | Why it mattered |
|---|---|---|
| Included full-text reports | 11 | Large enough to test more than a single clean PDF. |
| Configured protocol outcomes | 17 | Kept the agent tied to the review question rather than free-form outcome hunting. |
| Extraction target | Study-level outcome proposals | Matched how reviewers normally work through one included study at a time. |
| Write model | Proposal-only | Protected the canonical extraction sheet until researcher acceptance. |
What was tested
The agent was asked to work study-wise: review an included study, focus on the configured outcomes, prepare multiple outcome proposals, and submit those proposals for researcher review.
Coventra enforced the important boundaries: no final extraction row was written directly by the agent, missing values had to stay missing, and repeat submissions were checked before entering the queue.
- Protocol-guided extraction: the agent worked from configured outcomes rather than inventing endpoints.
- PDF grounding: proposed values had to be tied to the uploaded full text.
- Multi-outcome submission: one study could produce several outcome proposals in a single review pass.
- Completeness checks: incomplete outcome drafts were kept out of the reviewer queue.
- Repeat-run safety: a repeated agent run did not duplicate proposals already present in the project.
Results
The fresh extraction pass completed successfully. The agent worked from the included-study context, prepared study-level outcome proposals, and validated three outcome rows before final submission.
The final submission reached the expected pending-review state. When the same rows were submitted again, Coventra's repeat-run protection detected all three as repeats and prevented review-queue flooding.
The benchmark also confirmed that the public workflow is reviewer-safe: the agent can move work forward, but accepted extraction data still requires researcher action inside Coventra.
Fresh extraction benchmark results
| Check | Observed result | Public interpretation |
|---|---|---|
| PDF readiness | 11 included reports ready | The agent had enough full-text material to begin extraction work across the included evidence set. |
| Protocol coverage | 17 configured outcomes | Extraction was anchored to the review protocol, not ad-hoc endpoint selection. |
| Study-wise extraction pass | 25 outcome-review targets prepared for the first study-level pass | The workflow can give the agent a concrete reviewer-style task instead of asking it to wander through a PDF. |
| Validation | 3 outcome rows passed pre-submission validation | Values could be checked before entering the review queue. |
| Submission state | Pending review | Agent output remained reviewable and did not become canonical data automatically. |
| Duplicate protection | 3 repeated rows detected and not duplicated | Safe resume behavior: rerunning the agent does not flood the project with repeated suggestions. |
| Focused evidence | No oversized-response failures in the fresh pass | The agent received focused source evidence rather than a raw document dump. |
What this means for reviewers
The extraction benchmark shows the intended working model: Coventra narrows the extraction task, the AI proposes source-backed values, and the researcher accepts, edits, or rejects inside the sheet.
This is especially useful for outcomes that are easy to miss in long oncology reports: survival endpoints, response rates, adverse-event summaries, subgroup caveats, and table footnotes. The benchmark does not mean the AI should be trusted blindly; it means the workflow keeps the AI's work inspectable.
- The agent is useful for locating and drafting values, not for final scientific judgment.
- Configured outcomes reduce cherry-picking risk because the target list comes from the protocol.
- Source-backed proposals reduce verification time because the reviewer starts from the cited evidence.
- Repeat-run protection makes it practical to resume interrupted extraction sessions.
Limits and next validation
This benchmark is positive and launch-useful, but it is still a workflow benchmark. It verifies that source-backed suggestions, the review queue, provenance, validation checks, and duplicate protection work together.
The next layer is a locked answer-key benchmark: compare extracted values against a pre-specified human gold standard across baseline tables, safety tables, survival endpoints, repeated values, missing outcomes, and messy PDF layouts.
- Use MCP extraction suggestions as a fast reviewer-assist layer.
- Verify every accepted value against the source PDF before analysis.
- Do not treat this benchmark as a claim that every outcome in every PDF can be extracted without human correction.