AI-Assisted Evidence Synthesis & Research Evaluation
Structured research briefings with explicit verification status
AI-Assisted Research / Evidence Evaluation
Using AI to process research material while keeping critical evaluation of the evidence in human hands. Research articles were turned into structured briefings that separate findings from the evidence behind them, and label what is supported by the source versus what still requires verification.
01 — The Problem
A general research-summary prompt could produce a readable summary, but it did not explicitly require:
- Evidence attached to findings
- Source/section identification
- Evidence classification
- Verification status
- Separation between source statements and interpretation
This created a risk of blurring what the source actually said with what the AI inferred.
02 — Baseline Prompt
The original prompt used as a baseline:
Summarise the following research material. Identify the main findings, supporting evidence, limitations, and unanswered questions. Present the information clearly and concisely.03 — What Needed Improvement
- Insufficient evidence traceability
- No explicit verification status
- No evidence-type classification
- No explicit instruction against unsupported additions
- Limited separation between source claims and interpretation
04 — Prompt Iteration
The prompt was redesigned to require explicit evidence and verification controls: evidence attached to each finding, source/section identification, evidence type, verification status, an instruction against unsupported additions, and a clear separation between source statements and interpretation.
05 — Testing
The revised workflow (V2) was tested against two different research sources.
Test 1
Automation and AI in the Workplace: The Future of Work Is More Complex Than Ever (2024)
The workflow surfaced:
- Evidence traceability
- Source/section information
- Verification status
- Third-party statistics requiring verification
- Limitations in the supplied case study
- Missing methodology/raw data
Test 2
Streamlining Space Processes through IT Automation: Reducing Manual Tasks for Enhanced Efficiency in Space Exploration (2019)
The workflow surfaced:
- Unresolved placeholder citations
- Unsupported use of the word "significant"
- Missing quantitative data
- Missing experimental conditions
- Missing expert-interview details
- Inconsistencies in the supplied reference list
06 — V1 vs V2
| Criterion | V1 | V2 |
|---|---|---|
| Key findings | Yes | Yes |
| Clear structure | Yes | Yes |
| Evidence attached to findings | Limited | Yes |
| Evidence type identified | No | Yes |
| Source/section identified | No | Yes |
| Source vs interpretation | Limited | Yes |
| Verification status | No | Yes |
| Unsupported claims flagged | Limited | Yes |
| Limitations/unanswered questions | Yes | Yes, more rigorously |
07 — Results
- V2 produced evidence-linked findings rather than a summary alone.
- Third-party statistics and claims requiring independent verification were explicitly flagged.
- Author assertions were distinguished from independently verified evidence.
- The second test surfaced unresolved citations, missing quantitative support and incomplete methodological details.
- The same V2 workflow maintained its intended verification behaviour across two different research topics.
08 — Responsible AI
- AI was restricted to the supplied research material for the synthesis task.
- Unsupported information was labelled rather than invented.
- Cited sources were not automatically treated as independently verified.
- Verification status was made visible.
- Human review remained necessary for professional decision-making.
09 — Documentation
The prompt text, the evaluation criteria and the V1 versus V2 comparison are reproduced in full on this page. Source articles are third-party publications and are not redistributed here.
13 — What I learned
- How prompt structure — not prompt length — determines whether an AI output can be checked.
- How to classify evidence types and hold AI output to a verification standard.
- How to document research in a form another reader can audit.
14 — Limitations
- External statistics quoted inside the source articles were not independently verified against their original publications. The project's purpose was to identify what the source supported and what required verification, not to complete that verification.
- Two articles is a small sample; the prompt pattern is demonstrated, not benchmarked.
- This is AI-assisted research and evidence organisation. It is not original scientific research and not AI system development.
15 — Skills demonstrated
16 — Project evidence
Evidence for this project is the prompt text, the evaluation criteria and the comparison table reproduced in full on this page.