AI-Assisted Operations Request Triage & Workflow Evaluation
Prompt evaluation on a 20-request synthetic dataset
AI-Assisted Workflow / Operations
A structured evaluation of an AI-assisted triage workflow for incoming operational requests. Two prompt versions were tested against a synthetic 20-request dataset and scored on category accuracy, urgency accuracy, missing-information flagging, human-review escalation and overclaiming of completed actions.
- Synthetic / demonstration data — no real company or operational data is used anywhere in this project.
Workflow design
The objective was to design and evaluate an AI-assisted workflow that classifies workplace requests, extracts key information, flags missing details and risks, drafts responses, and routes consequential cases to human review. This is a prototype evaluation on synthetic data — not a production deployment, and it makes no claim of real-world time or cost savings.
01 — The Problem
An unconstrained triage prompt can classify requests and draft replies, but it can also quietly:
- Invent completed actions that never happened
- Over-rate urgency based on tone rather than explicit evidence
- Skip missing-information checks entirely
- Provide no mechanism to escalate sensitive actions to a human
In a real operations workflow this would create rework, compliance risk and potential security issues.
02 — Baseline Prompt
The original, unconstrained prompt used as the baseline:
Classify each workplace request by category and urgency. Summarise the request, say what action should be taken, and draft a professional response.03 — What Needed Improvement
- Free-text categories drifted away from the intended taxonomy
- Urgency was driven by phrasing and tone, not by explicit rules
- Missing information was never surfaced before a response was drafted
- Draft responses repeatedly asserted that actions were already complete
- No escalation path for credential changes, conflicting records or approvals
04 — Prompt Iteration
The prompt was redesigned with explicit constraints: a fixed category taxonomy, evidence-based urgency rules, mandatory missing-information flagging, a prohibition on inventing completed actions, and defined triggers for human review.
05 — The Dataset
A synthetic dataset of 20 workplace requests across Finance, HR, Operations, Procurement and IT. Each request was labelled with an expected category and urgency to create a controlled gold standard.
20
Synthetic requests
5
Departments represented
6
Fixed categories
Test coverage was deliberately mixed: clear requests, incomplete requests, ambiguous urgency, conflicting records, technical access problems, financial approvals, credential handling and explicit deadlines. No real employee, customer, supplier or company information was used.
06 — Testing
Both prompts were run once against the same 20 requests in a controlled comparison. V1 outputs were free-form; V2 outputs were structured against the new rules. Each output was scored against the gold standard for category accuracy, urgency accuracy, evidence-based reasoning, missing-information detection, risk detection, draft-response quality and human-review escalation.
This is a controlled comparison of prompt design, not a benchmark across models or repeated runs. The gold standard defines expected category and urgency; missing-information, risk and human-review values reflect evaluator judgement.
07 — V1 vs V2
| Metric | V1 | V2 |
|---|---|---|
| Category accuracy (of 20) | 20 | 20 |
| Urgency accuracy (of 20) | 16 | 20 |
| Missing information explicitly flagged | Never | Every request |
| Human-review escalation triggered | No mechanism | 5 of 20 |
| Draft responses overclaiming a completed action | 10 | 0 |
08 — Where V1 Failed
Category accuracy was identical at 20/20, but the practical gap was much larger. V1 failed on urgency in four requests and overclaimed completed actions in half the dataset.
Urgency failures (4 of 20)
| Request | Topic | V1 urgency | Why it was wrong |
|---|---|---|---|
| REQ-002 | Onboarding checklist | High | V1 rated High from tone ('starting Monday'); V2 correctly held Medium because no deadline for the checklist itself was stated. |
| REQ-008 | Client report | High | V1 treated 'ASAP' as High; V2 withheld High because no explicit deadline was given, matching the gold standard (Medium). |
| REQ-009 | PO quantity mismatch | Medium | V1 under-rated urgency and picked one of two conflicting quantities. V2 flagged HIGH and human review for conflicting records. |
| REQ-010 | Password reset to personal Gmail | Medium | V1 under-rated urgency and claimed the password 'has been sent' to a personal Gmail — enacting the security risk V2 was designed to catch. |
Overclaimed actions (10 of 20)
| Request | Topic | What V1 claimed |
|---|---|---|
| REQ-003 | Report issue | V1 claimed the report was already 'corrected'. |
| REQ-005 | Folder access lost | V1 claimed access 'has been restored' without confirmation. |
| REQ-007 | Spreadsheet review | V1 claimed to have reviewed a never-supplied spreadsheet. |
| REQ-011 | Directory update | V1 claimed the directory 'has been updated' without the new mapping. |
| REQ-013 | Supplier follow-up | V1 claimed the supplier 'has been followed up' without a named supplier. |
| REQ-014 | Dashboard refresh | V1 claimed the refresh 'ran successfully', inventing a system status. |
| REQ-015 | Payment approval | V1 stated the payment 'has been approved and processed' — a fabricated approval. |
| REQ-017 | Report update | V1 claimed the report was updated and 'sent to the customer' without the attachment. |
| REQ-019 | Portal outage | V1 claimed the outage 'should be back online shortly', implying an unconfirmed fix. |
09 — Human Review Triggers
V2 flagged 5 of 20 requests for human review — not because the AI failed, but because the rules identified decisions that should not be automated:
- REQ-006: New-starter account and credential provisioning
- REQ-009: Conflicting purchase-order and supplier-email quantities
- REQ-010: Password reset requested to a personal Gmail address
- REQ-015: Payment approval with a legal-threat escalation
- REQ-019: Full client-portal outage before a customer demo
10 — Results
- V2 matched all 20 expected urgency labels; V1 matched 16.
- V2 flagged missing information on every request instead of assuming details.
- V2 produced zero overclaimed actions, compared with 10 for V1.
- V2 escalated credential, conflict, approval, security and outage requests to human review.
- The structured constraints changed the practical usefulness of the output far more than the headline accuracy gap suggests.
11 — Responsible AI
- AI was restricted to the supplied request text and prohibited from inventing actions.
- Missing information was surfaced rather than silently filled in.
- Sensitive or consequential actions were routed to human review.
- Urgency was tied to explicit evidence, not tone or urgency words alone.
- The evaluation was transparent about its synthetic-data and single-run limitations.
12 — Documentation
The full project files — evaluation workbook, synthetic dataset, blank evaluation template and case-study draft — are available below.
13 — What I learned
- How to specify a workflow so an AI step is auditable rather than merely plausible.
- How to design an evaluation dataset that exposes failure modes instead of confirming success.
- How to structure a workbook so the flow of work is visible to a reader.
14 — Limitations
- The dataset is synthetic and written for testing. No real operational or company data was used.
- Twenty requests is a demonstration sample, not a statistical benchmark.
- This is practical digital operations work. No software was engineered and no AI model was built or trained.
15 — Skills demonstrated
16 — Project evidence
The evaluation workbook, synthetic dataset, evaluation template and case study draft are all downloadable on this page.