All projects
02CompletedPersonal / self-directed project

AI-Assisted Operations Request Triage & Workflow Evaluation

Prompt evaluation on a 20-request synthetic dataset

AI-Assisted Workflow / Operations

A structured evaluation of an AI-assisted triage workflow for incoming operational requests. Two prompt versions were tested against a synthetic 20-request dataset and scored on category accuracy, urgency accuracy, missing-information flagging, human-review escalation and overclaiming of completed actions.

  • Synthetic / demonstration data — no real company or operational data is used anywhere in this project.
Prompt engineeringMicrosoft ExcelSynthetic test datasetStructured evaluation criteria

Workflow design

Incoming requestAI classificationStructured extractionMissing-information checkDraft responseHuman reviewFinal action

The objective was to design and evaluate an AI-assisted workflow that classifies workplace requests, extracts key information, flags missing details and risks, drafts responses, and routes consequential cases to human review. This is a prototype evaluation on synthetic data — not a production deployment, and it makes no claim of real-world time or cost savings.

01 — The Problem

An unconstrained triage prompt can classify requests and draft replies, but it can also quietly:

  • Invent completed actions that never happened
  • Over-rate urgency based on tone rather than explicit evidence
  • Skip missing-information checks entirely
  • Provide no mechanism to escalate sensitive actions to a human

In a real operations workflow this would create rework, compliance risk and potential security issues.

02 — Baseline Prompt

The original, unconstrained prompt used as the baseline:

Classify each workplace request by category and urgency. Summarise the request, say what action should be taken, and draft a professional response.

03 — What Needed Improvement

  • Free-text categories drifted away from the intended taxonomy
  • Urgency was driven by phrasing and tone, not by explicit rules
  • Missing information was never surfaced before a response was drafted
  • Draft responses repeatedly asserted that actions were already complete
  • No escalation path for credential changes, conflicting records or approvals

04 — Prompt Iteration

The prompt was redesigned with explicit constraints: a fixed category taxonomy, evidence-based urgency rules, mandatory missing-information flagging, a prohibition on inventing completed actions, and defined triggers for human review.

05 — The Dataset

A synthetic dataset of 20 workplace requests across Finance, HR, Operations, Procurement and IT. Each request was labelled with an expected category and urgency to create a controlled gold standard.

20

Synthetic requests

5

Departments represented

6

Fixed categories

Test coverage was deliberately mixed: clear requests, incomplete requests, ambiguous urgency, conflicting records, technical access problems, financial approvals, credential handling and explicit deadlines. No real employee, customer, supplier or company information was used.

06 — Testing

Both prompts were run once against the same 20 requests in a controlled comparison. V1 outputs were free-form; V2 outputs were structured against the new rules. Each output was scored against the gold standard for category accuracy, urgency accuracy, evidence-based reasoning, missing-information detection, risk detection, draft-response quality and human-review escalation.

This is a controlled comparison of prompt design, not a benchmark across models or repeated runs. The gold standard defines expected category and urgency; missing-information, risk and human-review values reflect evaluator judgement.

07 — V1 vs V2

MetricV1V2
Category accuracy (of 20)2020
Urgency accuracy (of 20)1620
Missing information explicitly flaggedNeverEvery request
Human-review escalation triggeredNo mechanism5 of 20
Draft responses overclaiming a completed action100

08 — Where V1 Failed

Category accuracy was identical at 20/20, but the practical gap was much larger. V1 failed on urgency in four requests and overclaimed completed actions in half the dataset.

Urgency failures (4 of 20)

RequestTopicV1 urgencyWhy it was wrong
REQ-002Onboarding checklistHighV1 rated High from tone ('starting Monday'); V2 correctly held Medium because no deadline for the checklist itself was stated.
REQ-008Client reportHighV1 treated 'ASAP' as High; V2 withheld High because no explicit deadline was given, matching the gold standard (Medium).
REQ-009PO quantity mismatchMediumV1 under-rated urgency and picked one of two conflicting quantities. V2 flagged HIGH and human review for conflicting records.
REQ-010Password reset to personal GmailMediumV1 under-rated urgency and claimed the password 'has been sent' to a personal Gmail — enacting the security risk V2 was designed to catch.

Overclaimed actions (10 of 20)

RequestTopicWhat V1 claimed
REQ-003Report issueV1 claimed the report was already 'corrected'.
REQ-005Folder access lostV1 claimed access 'has been restored' without confirmation.
REQ-007Spreadsheet reviewV1 claimed to have reviewed a never-supplied spreadsheet.
REQ-011Directory updateV1 claimed the directory 'has been updated' without the new mapping.
REQ-013Supplier follow-upV1 claimed the supplier 'has been followed up' without a named supplier.
REQ-014Dashboard refreshV1 claimed the refresh 'ran successfully', inventing a system status.
REQ-015Payment approvalV1 stated the payment 'has been approved and processed' — a fabricated approval.
REQ-017Report updateV1 claimed the report was updated and 'sent to the customer' without the attachment.
REQ-019Portal outageV1 claimed the outage 'should be back online shortly', implying an unconfirmed fix.

09 — Human Review Triggers

V2 flagged 5 of 20 requests for human review — not because the AI failed, but because the rules identified decisions that should not be automated:

  • REQ-006: New-starter account and credential provisioning
  • REQ-009: Conflicting purchase-order and supplier-email quantities
  • REQ-010: Password reset requested to a personal Gmail address
  • REQ-015: Payment approval with a legal-threat escalation
  • REQ-019: Full client-portal outage before a customer demo

10 — Results

  • V2 matched all 20 expected urgency labels; V1 matched 16.
  • V2 flagged missing information on every request instead of assuming details.
  • V2 produced zero overclaimed actions, compared with 10 for V1.
  • V2 escalated credential, conflict, approval, security and outage requests to human review.
  • The structured constraints changed the practical usefulness of the output far more than the headline accuracy gap suggests.

11 — Responsible AI

  • AI was restricted to the supplied request text and prohibited from inventing actions.
  • Missing information was surfaced rather than silently filled in.
  • Sensitive or consequential actions were routed to human review.
  • Urgency was tied to explicit evidence, not tone or urgency words alone.
  • The evaluation was transparent about its synthetic-data and single-run limitations.

12 — Documentation

The full project files — evaluation workbook, synthetic dataset, blank evaluation template and case-study draft — are available below.

13 — What I learned

  • How to specify a workflow so an AI step is auditable rather than merely plausible.
  • How to design an evaluation dataset that exposes failure modes instead of confirming success.
  • How to structure a workbook so the flow of work is visible to a reader.

14 — Limitations

  • The dataset is synthetic and written for testing. No real operational or company data was used.
  • Twenty requests is a demonstration sample, not a statistical benchmark.
  • This is practical digital operations work. No software was engineered and no AI model was built or trained.

15 — Skills demonstrated

Structured data handlingData cleaning and organisationWorkflow logicSpreadsheet designAI-assisted analysisDocumentation

16 — Project evidence

The evaluation workbook, synthetic dataset, evaluation template and case study draft are all downloadable on this page.