AI-Enabled Regression Test Optimization
An open-source reference for using AI to decide which regression tests a change actually needs, and for automating those tests end to end with PyTest — on real hardware, not in a simulator.
Introduction
Testing every build in full is the safe choice and the slow one. On embedded products the cost is bench time: the device takes one connection at a time, tests run in real time, and a full suite can take longer than the gap between builds. The usual fix is to let an engineer pick a subset by hand — fast, but not repeatable, not recorded, and only as good as their memory of the change.
This project replaces that guesswork with an explicit pipeline: work out what changed, let an AI planner pick the smallest set of tests that would catch a regression, then generate, run and record them automatically with PyTest — on real hardware, not in a simulator.
Architecture
A target-independent core sits above the line; one adapter per product sits below it. The core reads the evidence, plans and generates the tests, and runs them — while everything specific to a given board lives in the replaceable adapter layer.

How the framework works
The framework turns test selection into four plain steps that run the same way every time.
- 1
Evidence
Work out what changed — from the release note, the built firmware image or the source — and read the current state of the device.
- 2
Decision
An AI planner picks the smallest set of tests, and how hard to push them, that would catch a regression in that change. It chooses only from tests the framework can actually run, and a simple rule-based planner is the default and the fallback.
- 3
Automation
The chosen tests are written as a PyTest module, checked mechanically, run against the device, and reported. Requirements and release notes are turned into tests automatically on every run.
- 4
Feedback
If a test fails, the intensity is raised and the run repeated. Every run is recorded, so the framework can show whether it is working.
Where the AI is used
The AI makes one decision: given this change and this device state, which tests are worth running, and how hard.
- ✓Choosing which tests are worth running for this change, and the stress intensity for each
- ✓Weighing free-form evidence — a release-note sentence, a risk score, metrics that may be missing
- ✓A schema-constrained planner, opt-in, with a deterministic rule-based fallback
Requirement and release-note tests
Not every test is chosen by the AI. On every run, the framework also reads the requirements list and the release notes and writes a PyTest case for each one, automatically.
Each requirement becomes a check that it holds — plus extra checks right at any limit it states. Each change in the release notes becomes the tests that change implies. The test always follows the requirement, so editing the requirement moves the test with it; a case that cannot run is skipped with a reason, never quietly dropped.
So every requirement stays tied to a test that checks it — kept in step by the framework, not by hand — and nothing is left untested.
Proof of concept: nRF52840
The framework is target-independent. To show that it works on real hardware, it is applied to a Nordic nRF52840 development board running Zephyr firmware — an audio compressor that stands in for a hearing aid — reached over Bluetooth Low Energy.
The board is the demonstration, not the subject: everything specific to it lives in one replaceable adapter layer, so porting to another target means rewriting that layer and nothing else.
Demonstration
AI-selected regression tests generated and run against the nRF52840 over Bluetooth Low Energy.
Engineering Demonstrated
- ✓Risk-based test selection (weighted threshold model)
- ✓AI planning layer with schema-constrained LLM and offline fallback
- ✓Dynamic pytest generation (AST-validated)
- ✓Failure-driven escalation and re-run under stress
- ✓Auto-generated requirement and release-note traceability tests
- ✓BLE / GATT transport to nRF52840 / nRF5340 (Zephyr, Cortex-M4F)
- ✓On-device WDRC audio DSP verification (SNR loopback)
- ✓JUnit XML / CI-compatible reporting
Where This Can Help
- ✓Embedded firmware regression testing on real hardware
- ✓Hearing-aid and audio DSP validation
- ✓Reducing regression execution time and effort
- ✓BLE device test automation
- ✓Risk-prioritised test suites in CI pipelines
- ✓Coverage-preserving test reduction with generated traceability
Open Source
Explore the source code and documentation for the AI-Enabled Regression Test Optimization project.
View Source Code