← Back to Work

Demonstration Study • AI & Workflow Automation

Automate the right work. Keep people in control.

A workflow opportunity assessment and a working inbox-screening model. Explore what is worth changing, what the evidence supports and where a person should make the final decision.

5,574 source messages763 held-out test messagesPublic UCI datasetIndependent demonstration

Context

A busy inbox is not a reason to automate every decision.

Start with the work, measure the errors, then design the handoff.

The business question is where AI can remove repetitive handling without losing a legitimate enquiry or creating another review burden. This demonstration separates an editable business case from a measured text-classification experiment.

The working example uses historical, publicly available SMS messages to screen incoming text for spam. It is a small machine-learning model, not a generative chatbot. It proposes a queue; a person can override that proposal and record the decision.

Demonstration note: This is not a client deployment. Test results describe this English SMS corpus only. Planning inputs are illustrative assumptions, and the source cannot establish productivity gains in a modern business inbox.

01 · Opportunity assessment

Make the business case visible before building.

Adjust the illustrative scenario below. Account for assisted handling, adoption and running cost, then check whether the workflow is ready for a bounded test.

Workflow and assumptions

Planning scenario · editable assumptions

Proposed rubric: readiness is the mean of the three ratings. An absent owner, weak data or high consequence changes the recommendation. This is a planning aid, not a certified assessment or automatic approval.

02 · Working automation

A proposal, a review, a visible decision.

The trained model runs in your browser. Choose a held-out public example or enter short English text. Your typed message is processed locally and is not sent to a server.

01

Receive

Take the incoming text as data. Do not execute instructions inside it.

02

Screen

Apply the trained model and a vocabulary-support check.

03

Review

Keep uncertain inputs in manual triage. Let a person override any route.

04

Record

Record the proposal and final decision. Do not automatically delete a message.

Try the screening workflow

Examples retain the source wording and published labels.

This historical SMS model has not been validated on business email, WhatsApp enquiries, Nigerian languages or current spam campaigns. Do not use this demo to make a production decision.

Loading the model…

Human review desk

Every route is a proposal. Nothing is connected to a live inbox.

The demonstration log stays in memory until this page closes. Its CSV excludes message content. Confirming a decision does not send, delete or update anything outside this page.

03 · Measured evidence

Accuracy is a start. The errors decide the workflow.

The model was trained on 3,560 messages and checked on a separate validation set before reporting the untouched 763-message test set. Normalized exact duplicates were removed before splitting. The test contains 675 legitimate messages and 88 spam messages.

97.4%Held-out accuracy at a 50% classification threshold
88.8%Held-out spam F1 at that threshold
88.5%Accuracy of an always-legitimate baseline

What the default classifier gets wrong

763 held-out messages · threshold 50%
Published labelPredicted legitimatePredicted spam
Legitimate66411
Spam979

Eleven legitimate messages were flagged as spam; nine spam messages were missed. That makes silent deletion an unsuitable default, even with high headline accuracy.

Test accuracy is within this corpus. Near-duplicate templates may remain; this random split does not measure future performance or transfer to another channel.

Change the review policy

Scores at or above this level go to spam review. Scores below 20% go to the normal queue. The remaining messages go to manual triage.

  • Normal queue 649
  • Manual triage 40
  • Spam review 74
  • Legitimate messages flagged for spam review 2
  • Spam entering the normal queue 6

This threshold experiment uses held-out model scores. The interactive text tool adds a vocabulary-support guard. Scores are not independently calibrated certainty. Changing the threshold here explores the test set; it must not be presented as independent tuning evidence.

04 · Operating design

Move from a useful example to a controlled implementation.

Map the real workflow

Identify the owner, inputs, handoffs and exception types. Measure actual handling time, rework and missed messages before setting a productivity target.

Validate on the intended channel

Build a permissioned, appropriately handled dataset of real business messages. Compare a simple rules baseline with the trained model and any proposed language-model alternative. Measure errors by language, message type and business consequence.

Run in shadow mode

Generate proposals without changing the live workflow. Review disagreements and record overrides. Set acceptable error rates and escalation rules before connecting any action.

Connect a bounded action

Start with reversible tagging or queue assignment. Keep approvals for outward messages and consequential actions. Monitor drift, review volume, incident rates and actual time released; retain a rollback path.

Method and reproducibility

A complete model, with its limits left visible.

The downloaded file contains 5,574 lines. After normalized-text deduplication, 5,086 messages remain. A seeded, stratified 70/15/15 split separates training, validation and test data. TF–IDF text features feed a logistic-regression classifier; only the training split fits the vocabulary and model. The same coefficients run in the browser and are checked against Python predictions.

Source: Almeida, T. & Hidalgo, J. (2011), SMS Spam Collection, UCI Machine Learning Repository, DOI 10.24432/C5CC84. UCI licenses the dataset under CC BY 4.0. ONE Light Analytics performed deduplication, modelling, evaluation and workflow design. The reported test metrics are model results, not client outcomes.

Have a repetitive information workflow?

Bring us the work. We will help identify what is worth changing.

ONE Light Analytics can assess the opportunity, test the model and design the handoffs before connecting automation to everyday operations.