Alimoudine Idrissou

Independent Researcher · Dakar, Senegal

Alimoudine Idrissou

AI Researcher in Mathematical Reasoning & LLM Evaluation

I study how language and multimodal models reason, where current evaluations break down, and how rigorous, verifiable benchmarks can reveal their limitations.

My work combines olympiad mathematics, proof auditing, benchmark design, scientific evaluation, and software engineering.

Illustrated portrait of Alimoudine Idrissou

Mathematical reasoning and proof auditing

Can a system tell an invalid step from one that is merely compressed?

Studying whether reasoning systems can produce, inspect, and repair mathematically valid arguments rather than merely plausible final answers. The interesting cases are not arithmetic slips but arguments that read well and quietly depend on something never established.

Evaluation validity and benchmark design

What has to be true of a task before a score on it means anything?

Designing original, hand-verified tasks and evaluation procedures that expose model limitations while controlling for ambiguity, contamination, and weak answer checking. Difficulty and validity are separate properties, and a hard benchmark can still measure the wrong thing.

Scientific and multimodal reasoning

When a task requires reading a figure and reasoning quantitatively, which part actually fails?

Evaluating how models integrate textual, visual, and quantitative evidence in scientific problem-solving settings, including tasks constructed so that neither the image nor the text alone is sufficient.

Multilingual and low-resource AI

What does rigorous evaluation look like for languages that mainstream benchmarks barely cover?

Building and evaluating systems for African languages and contexts that are often poorly represented in mainstream AI benchmarks, where the gap is both in model quality and in the measurement apparatus itself.

Research statement and methodology


Work in progress2026 –

PremiseGuardDetecting Unsupported Load-Bearing Claims in Mathematical Proofs

Research question

Can language-model-based mathematical critics reliably detect, localize, explain, and repair unsupported load-bearing claims while distinguishing them from ordinary mathematical errors and acceptable omitted details?

PremiseGuard investigates a failure mode in which a proof silently relies on a nontrivial claim that has not been established under the stated assumptions. The project studies controlled proof variants, earliest-error localization, missing-premise reconstruction, dependency analysis, and proof repair.

Read the full description


Work in progressResearch and evaluation2026 –

PremiseGuard

Detecting Unsupported Load-Bearing Claims in Mathematical Proofs

Problem
A proof can silently depend on a nontrivial claim that was never established under the stated assumptions, and read as entirely convincing.
Contribution
An annotation scheme and task definition that separate unsupported load-bearing claims from ordinary errors and from acceptable compressed reasoning.
Role
Independent researcher
Proof auditingBenchmark designEvaluation methodology
CompletedMultilingual and educational AI2023

AfroXLMR NER

Named entity recognition for 21 African languages

Problem
Named entity recognition is well served for major languages and largely unavailable as usable tooling for African ones.
Contribution
A Python application that runs entity extraction over text and PDF input for 21 African languages, built on the Masakhane AfroXLMR NER model.
Role
Co-author — application development and evaluation · with Alain Ogou
PythonTransformersNERAfrican languages
CompletedMultilingual and educational AI2025

GandaLab / SymLab

An interactive mathematics and science learning environment

Problem
Abstract mathematical and scientific concepts are difficult to build intuition for when students never get to vary them and watch what happens.
Contribution
An interactive learning environment pairing video and simulation with an AI tutor that answers from the material and generates comprehension checks tied to specific source timestamps.
Role
Co-creator and lead developer · Team project — hackathon build
EducationInteractive learningAI tutoring
CompletedMultilingual and educational AI2025

Hakili

A French-language digital literacy platform

Problem
Digital literacy material on misinformation, deepfakes, and online safety is scarce in French for West African audiences.
Contribution
A conversational learning platform combining structured French course modules with an assistant that discusses a submitted link's credibility rather than pronouncing a verdict on it.
Role
Full-stack engineer · Team project
Next.jsConversational AIDigital literacy

All projects


5 min read

Building a FastAPI Starter

In this blog post, I share my journey in building a FastAPI auth starter code, highlighting the common dilemma software engineers face: whether to use an existing library or build a custom solution. While libraries like fast-sso offer quick and reliable solutions, they may not always fit the unique needs of every project. Recognizing this, I created a FastAPI auth starter that provides the best of both worlds—quick setup and flexibility. This starter code is designed to help developers save time and adapt to various types of applications, from small projects to complex SaaS products. Whether you need basic authentication or more advanced roles and permissions, this starter code can give you a head start. The post emphasizes the importance of choosing the right tool for your project, team, and goals, and encourages developers to consider both libraries and custom solutions based on their specific needs.

  • FastAPI

4 min read

Watching YouTube Tutorials as a Developer

This blog shares my experience of learning from YouTube tutorials. I discovered that instead of following every step, focusing on the tools, understanding best practices, and solving challenges helped me grow faster as a developer. It’s not about copying tutorials, but applying the knowledge to real projects.

  • Growthmindset

7 min read

The Machine Learning Life cycle

This blog dives deep into the Machine Learning Life Cycle, outlining the essential stages from problem definition to model deployment and monitoring. It emphasizes the importance of clear problem identification, high-quality data collection, and robust model training. Each stage is critical to building scalable and reliable ML systems, with insights into selecting the right algorithms, avoiding common pitfalls like overfitting, and ensuring long-term model performance through continuous monitoring.

  • ML

All writing


Alimoudine Idrissou is an independent AI researcher and software engineer based in Dakar, Senegal. His work focuses on mathematical reasoning, language-model evaluation, scientific multimodal benchmarks, and reliable evaluation methodology. He has designed and hand-verified original olympiad-level mathematics problems and scientific evaluation tasks for frontier AI systems. His background combines mathematics, computer science, software engineering, and mathematics education. He holds an M.S. in Computer Science from ESMT Dakar and a B.S. in Mathematics and Computer Science from Université d'Abomey-Calavi.

Full biography, experience, and education

Research correspondence

I read messages about evaluation methodology, mathematical reasoning, and benchmark design.

alimoudineidrissou@gmail.com