Benchmark question
This benchmark asked whether an MCP-connected agent could screen a full title/abstract queue inside a real Coventra project, surface the relevant trial family, and advance to full-text decisions—while keeping all agent output in the reviewable agent lane and leaving the canonical record and PRISMA counts under researcher control.
The benchmark project was MCP SRMA Benchmark — First-line PD-(L)1 + Chemotherapy in Advanced nsqNSCLC. It was designed as a workflow-readiness and relevance test, not as a formal accuracy claim against an independent gold standard.
This article reports workflow behavior and system safeguards. It does not claim that agent screening is clinically definitive, and it should not be read as medical advice.
Project setup
The project was deliberately initialized without decisions or downstream review data so the MCP run would start from a clean state.
Benchmark starting state
| Item | Value | Purpose |
|---|---|---|
| Imported records | 421 bibliographic records with abstracts | A realistic screening queue rather than a toy dataset. |
| Abstract coverage | 421 of 421 records had abstracts | Allowed title/abstract screening without title-only fallbacks. |
| Protocol | Locked PICO with 8 inclusion and 7 exclusion criteria | Kept screening decisions tied to explicit criteria. |
| Outcomes | 17 configured protocol outcomes | Prepared the project for downstream extraction testing. |
| Quality template | Generic randomized-trial RoB template with 5 domains | Prepared the project for later quality-assessment testing. |
| Existing decisions | None | Prevented leakage from prior human or agent decisions. |
| Existing extraction / RoB / analysis | None | Made downstream automation gaps visible. |
Method
Claude Sonnet 4.6 connected to the benchmark project through Coventra MCP. The agent worked from the project itself rather than copied spreadsheets or pasted abstracts.
The run began with a project orientation step. Screening then proceeded in controlled batches, using standardized reasons for obvious exclusions and preserving the reviewer's ability to inspect the output.
Agent decisions were stored in the dedicated agent lane. They did not become canonical screening decisions, did not affect PRISMA counts, and remained available for human review and merge.
- Orient the agent on the connected project state.
- Work through the title/abstract queue in controlled batches.
- Use standardized exclusion categories for obvious irrelevant records where possible.
- Store results in the dedicated agent lane.
- Resume from remaining undecided records; already-decided records are excluded from later batches.
- Move likely or unresolved records toward full-text/PDF readiness after title/abstract screening.
Results
Claude Sonnet 4.6 completed title/abstract screening across all 421 records. The agent excluded 380 records and flagged 41 as potentially relevant for full-text review. In the first full-text pass, it identified 11 includable trial reports while separating companion publications, subgroup analyses, duplicate records, and secondary reports into exclusions.
All decisions were written through the agent lane and subsequently merged under researcher control, preserving the human review checkpoint and a PRISMA-safe audit trail. After merge, the project shows 11 included records, 19 eligible or pending further full-text and extraction work, and 391 excluded.
The relevance signal was strong. The included and potentially included records fell within the intended evidence family—first-line PD-(L)1 plus chemotherapy RCTs in advanced non-squamous NSCLC. Trial reports surfaced in the eligible set included KEYNOTE-189, RATIONALE-304, CameL, TASUKI-52, POSEIDON, CheckMate 9LA, IMpower, ORIENT-11, CHOICE-01, and similar primary reports from the target class.
- Full queue screened: all 421 records processed, none skipped due to missing abstracts.
- Relevant trial family correctly surfaced: included records matched the target evidence class without requiring manual triage of the full queue.
- Aggressive irrelevant filtering: reviews, meta-analyses, economic analyses, preclinical studies, and wrong-population records excluded with standardised reason codes.
- Subgroup, companion, and secondary publications handled at full text: separated from primary trial reports rather than included alongside them.
- Agent decisions stayed reviewable: no canonical screening rows were written without explicit researcher merge.
Screening outcomes
| Stage | Decision | Count |
|---|---|---|
| Title/abstract | Potentially relevant (Include / flag) | 41 |
| Title/abstract | Exclude | 380 |
| Full-text (agent) | Include | 11 |
| Full-text (agent) | Exclude | 11 |
| After researcher merge | Included | 11 |
| After researcher merge | Eligible / pending further review | 19 |
| After researcher merge | Excluded | 391 |
These results were not compared against an independent locked answer key. Sensitivity, specificity, and recall against a gold standard are not claimed here. Accuracy validation requires a separate prospective benchmark with a pre-registered answer key.
Extraction follow-up
After screening and full-text progression, the same benchmark project was used for a separate MCP extraction benchmark. That follow-up tested whether the agent could work from included PDFs, target configured protocol outcomes, submit source-backed proposals, and keep the researcher review checkpoint intact.
The extraction results are now reported in a dedicated MCP extraction benchmark article so this screening report can stay focused on title/abstract and full-text progression.
- Read the MCP extraction benchmark for the PDF-grounded extraction results.
- The same propose-review-accept boundary applied in both screening and extraction.
The extraction follow-up is a workflow benchmark, not a complete clinical accuracy benchmark.
What the workflow demonstrated
The run demonstrated that an MCP-connected agent can complete the full title/abstract screening queue for a realistic oncology SRMA, advance to full-text decisions, and create PDF-grounded extraction proposals—inside Coventra's review model, without exporting the project to a separate tool.
The batch-and-resume workflow handled the 421-record queue without reprocessing already-decided records. The run also showed that the same project can move naturally from screening into PDF-grounded extraction while keeping human checkpoints intact.
- The MCP connector operated on live project data: criteria, study records, and decision state were read directly, not copied.
- Project-scoped consent and the agent lane kept the automation auditable and separated from the human reviewer record.
- The workflow reduced repetitive manual overhead while preserving reviewability.
- Extraction proposals remained read-only until accepted.
- Validation guards prevented incomplete extraction suggestions from entering the public workflow as if they were usable data.
- Human merge checkpoints remained intact: the agent suggested work; researchers controlled what became canonical.
Safety boundaries that held
- No canonical screening rows were written by the agent.
- Agent decisions stayed out of PRISMA and agreement calculations until explicitly merged.
- No canonical extraction rows were written by the agent; extracted values entered a pending review queue.
- Extraction proposals carried source provenance before they appeared in the UI.
- Project access was scoped to the connected project through per-project MCP consent, not a broad account-wide key.
- Audit logging avoided storing abstracts, private source quotes, credentials, or private prompts.
Limitations and next validation
This benchmark demonstrates end-to-end workflow readiness and a strong relevance signal, but it does not yet establish accuracy against an independent gold standard. The included records look correct for the target evidence family, but without a pre-registered answer key, recall and precision cannot be formally reported.
The extraction follow-up is reported separately. It verifies that PDF-grounded proposals, repeat-run protection, human acceptance, and validation checks can work together, but it does not yet validate every configured outcome, every adverse-event table shape, every baseline table, or every risk-of-bias domain.
The multi-PDF extraction session was interrupted by the external AI client's weekly usage limit, so the reported extraction numbers are the audited rows submitted before that limit was reached. The partial interruption is useful operationally—it shows the workflow can resume from existing proposals—but it is not a substitute for a full locked extraction benchmark.
The next validation layer is answer-key comparison: sensitivity and specificity for title/abstract screening, per-cell extraction accuracy with provenance verification, RoB domain agreement, and whether downstream analysis reproduces a published pooled estimate within the expected confidence interval.
- Use this benchmark as a public workflow and relevance report, not a substitute for protocol-specific human screening judgment.
- For formal accuracy claims, run a prospective benchmark with a locked answer key and publish sensitivity, specificity, and recall alongside the method definitions.