research
The Model Wrote the Anomaly Score Instead of Learning It
An in-context pipeline turned normal-state summaries into executable scoring logic and beat tested baselines across 24 datasets.

Summary
An in-context pipeline turned normal-state summaries into executable scoring logic and beat tested baselines across 24 datasets.
LLM-Detector summarizes normal tabular data into statistics, causal dependencies and prototypes, then asks a language model to generate a scoring engine for deviation, structural inconsistency and density. The authors compare it with 15 baselines across 24 mixed and continuous datasets and report consistent gains without model fine-tuning. The evidence is benchmark-based; generated scoring code still needs review, security controls and domain validation.
Why it matters
An in-context pipeline turned normal-state summaries into executable scoring logic and beat tested baselines across 24 datasets.
Limits and context
No additional limitation was separately recorded.
Key claims
An in-context pipeline turned normal-state summaries into executable scoring logic and beat tested baselines across 24 datasets.
Evidence: source-2026-08-21-008
Sources
- arXiv preprint 2608.19463arXiv · primary research
Corrections
No corrections have been recorded for this story.