research
The Classifier Gave Up Its Decision Rules
J-Miner turned internal signals into named concepts and compact executable rules that reproduced up to 98.3 percent of model decisions.
Summary
J-Miner turned internal signals into named concepts and compact executable rules that reproduced up to 98.3 percent of model decisions.
J-Miner aggregates vocabulary-aligned signals across layers and token positions, then learns explicit rules from the classifier’s own outputs. Across the reported tasks, those rules were 6.0 to 29.5 percentage points more faithful than equally compact word-based rules. Lightweight students with about one twenty-fourth the parameters retained 99.8 percent of mean source accuracy, showing a route to inspectable reuse without proving that every mined concept is causally decisive.
Why it matters
J-Miner turned internal signals into named concepts and compact executable rules that reproduced up to 98.3 percent of model decisions.
Limits and context
No additional limitation was separately recorded.
Key claims
J-Miner turned internal signals into named concepts and compact executable rules that reproduced up to 98.3 percent of model decisions.
Evidence: source-2026-08-19-005
Sources
- arXiv preprint 2608.17063arXiv · primary research
Corrections
No corrections have been recorded for this story.