# Project 04 — AI & Workflow Automation

A complete, independent capability demonstration by ONE Light Analytics: workflow opportunity assessment, trained inbox-screening model, measured errors, human review and an implementation plan.

## Source and licence

Almeida, T. & Hidalgo, J. (2011). SMS Spam Collection [Dataset]. UCI Machine Learning Repository. DOI: https://doi.org/10.24432/C5CC84.

Official dataset: https://archive.ics.uci.edu/dataset/228/sms+spam+collection

Official download: https://archive.ics.uci.edu/static/public/228/sms+spam+collection.zip

UCI identifies the dataset licence as CC BY 4.0: https://creativecommons.org/licenses/by/4.0/. Retain the source attribution when using this derivative. Modifications: normalized-text deduplication, split assignment, TF–IDF modelling, prediction and evaluation by ONE Light Analytics. `dataset.csv` contains the retained source messages and their split assignments. Its text is original source content, not newly invented business messages.

## Reproduce

Use Python 3.11+ and the pinned scikit-learn environment in `requirements.txt`.

```sh
python -m pip install -r projects/ai-workflow/requirements.txt
python projects/ai-workflow/build.py
node projects/ai-workflow/test-model.mjs
```

To reproduce from an already downloaded official archive:

```sh
python projects/ai-workflow/build.py --source-zip /absolute/path/to/sms-source.zip
```

The source hash and exact controls are in `validation.json`. No API key or paid AI service is needed. A static web server serves the case study and its same-origin JSON files. The model runs locally in the browser; input messages are not submitted to an API or persisted by the application.

## Analytical design

The downloaded file has 5,574 nonempty labelled lines. Tokenize lowercase text with ASCII `[a-z][a-z0-9]+`, normalize to a token sequence, remove duplicates before splitting, and exclude any conflicting-label groups. This retains 5,086 messages. No target label is passed as a feature.

Seed 42 produces a stratified 70/15/15 split: 3,560 training, 763 validation, 763 held-out test records. Fit a 2,500-word vocabulary and IDF only on training data. Use sublinear term frequency, L2 normalization and a logistic-regression classifier (`C=8`, balanced class weights). Validation results are reported separately; the fixed model configuration and threshold policy were not selected by optimizing held-out test results.

`model.json` exports full-precision coefficients and vocabulary. JavaScript scoring mirrors Python. Every held-out prediction is numerically checked within 1e-12 by `test-model.mjs`. The test verifies model parity, not an independent scientific validation of the corpus.

At the standard 0.50 classification cutoff: test accuracy 97.38%, spam precision 87.78%, recall 89.77%, F1 88.76%. The always-legitimate baseline accuracy is 88.47%. Confusion matrix (actual rows legitimate/spam; predicted columns legitimate/spam): [[664,11],[9,79]].

## Operational policy

Default illustrative score policy: score >=0.80 -> spam review; score <0.20 -> normal queue; otherwise manual triage. On the held-out score-only experiment this gives 74 spam-review, 649 normal, 40 manual; two legitimate messages enter spam review and six spam messages enter normal. Nothing is automatically deleted. The live tool additionally requires at least five tokens, three recognised token occurrences and 25% vocabulary coverage, otherwise manual triage. This guard has not itself been tuned or validated as an out-of-domain detector.

Slider changes are exploratory test-set analysis, not independent model-selection evidence. Model probabilities are not independently calibrated certainty. Review logs contain time, proposed route, score, text-support status and reviewer decision; message content is excluded. Logs exist in page memory until exported or the page closes.

## Workflow assessment

All economic and effort inputs are editable planning assumptions. Recoverable hours = monthly volume × (manual minutes − assisted minutes) × assisted share / 60. Assisted minutes must include checking and rework. Capacity value = hours × value per hour. Monthly value after running cost = capacity value − monthly cost. First-year value subtracts 12 monthly running costs and the setup cost. Simple payback divides setup cost by positive monthly net value; otherwise no positive payback is shown. No productivity percentage is inferred from model accuracy.

Readiness is the mean of data readiness, task repeatability and integration ease, each 1–5. Missing ownership, weak data and high-consequence decisions override the basic recommendation. This proposed rubric is a planning aid, not an industry standard or automatic deployment approval.

## Limits

Historical SMS from varied research collections is not a modern business inbox or Nigerian channel benchmark. Normalized exact duplicates are removed, but near-duplicates/templates may remain. Random within-corpus splitting does not measure future or cross-domain generalization. Published labels describe spam/legitimate text, not urgency, customer value or service eligibility. No measured client productivity, revenue increase or achieved financial saving is claimed. Native integrations, production monitoring and target-channel validation are future client-specific implementation work.
