AI Copilot

MCP benchmark with Claude Sonnet 4.6

Methods and workflow results from the first Coventra MCP benchmark run: title/abstract screening, full-text progression, and PDF-grounded extraction proposals in a real oncology SRMA project.

12 min readUpdated June 26, 2026

Benchmark question

This benchmark asked whether an MCP-connected agent could screen a full title/abstract queue inside a real Coventra project, surface the relevant trial family, and advance to full-text decisions—while keeping all agent output in the reviewable agent lane and leaving the canonical record and PRISMA counts under researcher control.

The benchmark project was MCP SRMA Benchmark — First-line PD-(L)1 + Chemotherapy in Advanced nsqNSCLC. It was designed as a workflow-readiness and relevance test, not as a formal accuracy claim against an independent gold standard.

Important

This article reports workflow behavior and system safeguards. It does not claim that agent screening is clinically definitive, and it should not be read as medical advice.

Project setup

The project was deliberately initialized without decisions or downstream review data so the MCP run would start from a clean state.

Benchmark starting state

ItemValuePurpose
Imported records421 bibliographic records with abstractsA realistic screening queue rather than a toy dataset.
Abstract coverage421 of 421 records had abstractsAllowed title/abstract screening without title-only fallbacks.
ProtocolLocked PICO with 8 inclusion and 7 exclusion criteriaKept screening decisions tied to explicit criteria.
Outcomes17 configured protocol outcomesPrepared the project for downstream extraction testing.
Quality templateGeneric randomized-trial RoB template with 5 domainsPrepared the project for later quality-assessment testing.
Existing decisionsNonePrevented leakage from prior human or agent decisions.
Existing extraction / RoB / analysisNoneMade downstream automation gaps visible.

Method

Claude Sonnet 4.6 connected to the benchmark project through Coventra MCP. The agent worked from the project itself rather than copied spreadsheets or pasted abstracts.

The run began with a project orientation step. Screening then proceeded in controlled batches, using standardized reasons for obvious exclusions and preserving the reviewer's ability to inspect the output.

Agent decisions were stored in the dedicated agent lane. They did not become canonical screening decisions, did not affect PRISMA counts, and remained available for human review and merge.

  1. Orient the agent on the connected project state.
  2. Work through the title/abstract queue in controlled batches.
  3. Use standardized exclusion categories for obvious irrelevant records where possible.
  4. Store results in the dedicated agent lane.
  5. Resume from remaining undecided records; already-decided records are excluded from later batches.
  6. Move likely or unresolved records toward full-text/PDF readiness after title/abstract screening.

Results

Claude Sonnet 4.6 completed title/abstract screening across all 421 records. The agent excluded 380 records and flagged 41 as potentially relevant for full-text review. In the first full-text pass, it identified 11 includable trial reports while separating companion publications, subgroup analyses, duplicate records, and secondary reports into exclusions.

All decisions were written through the agent lane and subsequently merged under researcher control, preserving the human review checkpoint and a PRISMA-safe audit trail. After merge, the project shows 11 included records, 19 eligible or pending further full-text and extraction work, and 391 excluded.

The relevance signal was strong. The included and potentially included records fell within the intended evidence family—first-line PD-(L)1 plus chemotherapy RCTs in advanced non-squamous NSCLC. Trial reports surfaced in the eligible set included KEYNOTE-189, RATIONALE-304, CameL, TASUKI-52, POSEIDON, CheckMate 9LA, IMpower, ORIENT-11, CHOICE-01, and similar primary reports from the target class.

  • Full queue screened: all 421 records processed, none skipped due to missing abstracts.
  • Relevant trial family correctly surfaced: included records matched the target evidence class without requiring manual triage of the full queue.
  • Aggressive irrelevant filtering: reviews, meta-analyses, economic analyses, preclinical studies, and wrong-population records excluded with standardised reason codes.
  • Subgroup, companion, and secondary publications handled at full text: separated from primary trial reports rather than included alongside them.
  • Agent decisions stayed reviewable: no canonical screening rows were written without explicit researcher merge.

Screening outcomes

StageDecisionCount
Title/abstractPotentially relevant (Include / flag)41
Title/abstractExclude380
Full-text (agent)Include11
Full-text (agent)Exclude11
After researcher mergeIncluded11
After researcher mergeEligible / pending further review19
After researcher mergeExcluded391
Important

These results were not compared against an independent locked answer key. Sensitivity, specificity, and recall against a gold standard are not claimed here. Accuracy validation requires a separate prospective benchmark with a pre-registered answer key.

Extraction follow-up

After screening and full-text progression, the same benchmark project was used for a separate MCP extraction benchmark. That follow-up tested whether the agent could work from included PDFs, target configured protocol outcomes, submit source-backed proposals, and keep the researcher review checkpoint intact.

The extraction results are now reported in a dedicated MCP extraction benchmark article so this screening report can stay focused on title/abstract and full-text progression.

Practical tips
  • Read the MCP extraction benchmark for the PDF-grounded extraction results.
  • The same propose-review-accept boundary applied in both screening and extraction.
Important

The extraction follow-up is a workflow benchmark, not a complete clinical accuracy benchmark.

What the workflow demonstrated

The run demonstrated that an MCP-connected agent can complete the full title/abstract screening queue for a realistic oncology SRMA, advance to full-text decisions, and create PDF-grounded extraction proposals—inside Coventra's review model, without exporting the project to a separate tool.

The batch-and-resume workflow handled the 421-record queue without reprocessing already-decided records. The run also showed that the same project can move naturally from screening into PDF-grounded extraction while keeping human checkpoints intact.

  • The MCP connector operated on live project data: criteria, study records, and decision state were read directly, not copied.
  • Project-scoped consent and the agent lane kept the automation auditable and separated from the human reviewer record.
  • The workflow reduced repetitive manual overhead while preserving reviewability.
  • Extraction proposals remained read-only until accepted.
  • Validation guards prevented incomplete extraction suggestions from entering the public workflow as if they were usable data.
  • Human merge checkpoints remained intact: the agent suggested work; researchers controlled what became canonical.

Safety boundaries that held

  • No canonical screening rows were written by the agent.
  • Agent decisions stayed out of PRISMA and agreement calculations until explicitly merged.
  • No canonical extraction rows were written by the agent; extracted values entered a pending review queue.
  • Extraction proposals carried source provenance before they appeared in the UI.
  • Project access was scoped to the connected project through per-project MCP consent, not a broad account-wide key.
  • Audit logging avoided storing abstracts, private source quotes, credentials, or private prompts.

Limitations and next validation

This benchmark demonstrates end-to-end workflow readiness and a strong relevance signal, but it does not yet establish accuracy against an independent gold standard. The included records look correct for the target evidence family, but without a pre-registered answer key, recall and precision cannot be formally reported.

The extraction follow-up is reported separately. It verifies that PDF-grounded proposals, repeat-run protection, human acceptance, and validation checks can work together, but it does not yet validate every configured outcome, every adverse-event table shape, every baseline table, or every risk-of-bias domain.

The multi-PDF extraction session was interrupted by the external AI client's weekly usage limit, so the reported extraction numbers are the audited rows submitted before that limit was reached. The partial interruption is useful operationally—it shows the workflow can resume from existing proposals—but it is not a substitute for a full locked extraction benchmark.

The next validation layer is answer-key comparison: sensitivity and specificity for title/abstract screening, per-cell extraction accuracy with provenance verification, RoB domain agreement, and whether downstream analysis reproduces a published pooled estimate within the expected confidence interval.

Practical tips
  • Use this benchmark as a public workflow and relevance report, not a substitute for protocol-specific human screening judgment.
  • For formal accuracy claims, run a prospective benchmark with a locked answer key and publish sensitivity, specificity, and recall alongside the method definitions.