Projects/ Research and evaluation
PremiseGuard
Detecting Unsupported Load-Bearing Claims in Mathematical Proofs
- Year
- 2026 –
- Role
- Independent researcher
- Team
- Solo
- Domain
- Mathematical reasoning · Evaluation methodology
PremiseGuard studies whether language-model-based mathematical critics can detect, localize, explain, and repair unsupported load-bearing claims in proofs. The project is in progress: task definitions, the annotation scheme, and the construction protocol are being developed.
Problem
A proof can silently depend on a nontrivial claim that was never established under the stated assumptions, and read as entirely convincing.
My contribution
An annotation scheme and task definition that separate unsupported load-bearing claims from ordinary errors and from acceptable compressed reasoning.
Role: Independent researcher.
Method
The design centres on controlled minimal pairs — a clean proof and a variant differing in one identifiable respect — together with explicit proof contracts that state the assumptions and permitted results, so that 'unsupported' has a definite meaning at each step. Premise attribution graphs record what each step depends on, which makes 'load-bearing' checkable rather than impressionistic. Hard negatives containing valid but compressed reasoning are included so that a critic is penalised for treating brevity as a defect.
Evaluation
No experiments have been run and no results are being claimed. The evaluation design under development covers four tasks — detection, earliest-error localization, explanation, and repair — scored separately, on the view that a single accuracy figure would obscure the distinctions the project exists to study.
Outcomes
- Task definitions for detection, localization, explanation, and repair
- An annotation scheme for load-bearing claims and proof contracts
- A construction protocol for controlled proof variants and hard negatives
Limitations
The scope is currently limited to proofs whose assumptions can be stated explicitly, which excludes much of ordinary mathematical writing. Judging whether an omitted step is acceptable depends on an assumed reader, and that judgement has to be pinned down before annotation can be consistent. No dataset, experimental result, or venue submission exists.
Related work
- AfroXLMR NER
Named entity recognition for 21 African languages
- GandaLab / SymLab
An interactive mathematics and science learning environment
- Hakili
A French-language digital literacy platform