Assemble domain-specific AI from a library of frozen specialist modules. Each module is independently trained. The system automatically selects the optimal modules for each input position. Add capabilities by adding modules. No retraining.
Modulith
Composable AI. Provable Control.
We build AI systems from independently trained frozen specialist modules — enabling provable capability control, full attribution, and regulatory compliance by construction.
Remove a capability by removing its module. Add it back — it returns. Instantly. No retraining, no fine-tuning, no gradient computation. Full per-position attribution trail shows exactly which module contributed what.
Mathematically verify that a capability is present or absent. Formal proofs in Lean 4. Designed for EU AI Act compliance, GDPR data removal, and regulatory frameworks requiring demonstrable AI governance.
Research
Formal Verification of Modular Compositional AI Systems
We investigate six safety properties that become structurally tractable in AI systems built from independently trained, permanently frozen, sparsely composed modules: attribution, guaranteed forgetting, non-degradation, cross-user isolation, bounded compute, and compositional monotonicity. We present empirical evidence from a production-scale system and outline a formal proof programme targeting machine-verified certificates in Lean 4.
Schulte-Zurhausen, K.D. (2026). "Open Problems in Formal Verification of Modular Compositional AI Systems." Modulith Research CIC. Accepted and presented at IEEE AIxSoftware 2026 (Summer Edition), Nagano; camera-ready in preparation.
The assessment methodology is deployed in production at Modulith Lab, where the same composition architecture evaluates frontier AI models against 24 formal properties.
Every output position decomposes exactly into per-module contributions via the composition graph. Routing weights provide a complete attribution trail at inference time.
Removing a module provably eliminates all associated knowledge. Empirically validated: removal residual below 1% of pre-removal capability. Output returns to exact baseline.
Adding modules cannot corrupt existing capabilities. Empirically validated: degradation bound below 2% across all tested configurations. Zero regression on module addition.
Inference cost is O(K) regardless of total pool size. A system with 100 modules and one with 10,000 modules run at identical speed. Formally proven.
User-specific modules have zero influence on other users. Architectural guarantee from independent training and frozen weights — no shared gradient, no information leakage.
Remove any module — output returns to exact baseline. Measured difference: 0.000000. Impossible in monolithic systems where knowledge, capability, and style are entangled.
Every compliance predicate in the Modulith assessment framework is defined as a Lean 4 formal specification. Lean 4 is an interactive theorem prover developed at Microsoft Research. When it accepts a proof, the result has been mechanically checked by a trusted kernel. No human judgment in the verification step. The following four theorem categories guarantee the integrity of every published assessment.
Every compliance predicate can be mechanically evaluated. Given a specification and a measurement, the answer is always PASS or FAIL — never undefined, never ambiguous. This eliminates the class of error where two evaluators disagree because the scoring rule is vague.
If a model passes a stricter threshold, it automatically passes any looser threshold. This prevents the paradox where tightening a standard could flip a result from FAIL to PASS — a class of bug that is impossible to detect without formal verification.
The overall assessment decomposes into independent sub-assessments. Passing Article 15 means passing accuracy AND robustness AND cybersecurity — this conjunction is proven, not assumed. Each sub-check is independently verifiable.
A model that abstains on uncertain queries (“I don’t know”) cannot lower its measured accuracy by doing so. Honest uncertainty is never penalised — a critical property for deployment in regulated contexts.
Full predicate definitions and theorem statements are published in the formal specifications repository. Proof implementations are available under commercial license. The Lean 4 kernel independently verifies every assessment — if the Python scoring and the Lean verification disagree on any property, the report is blocked from publication.
These proofs power every assessment published by Modulith Lab — 18 AI models assessed weekly against 24 Lean 4 verified properties. Reports issued independently by Modulith Research CIC.
Schulte-Zurhausen, K.D. (2026). Open Problems in Formal Verification of Modular Compositional AI Systems. Modulith Research CIC. Accepted and presented at IEEE AIxSoftware 2026 (Summer Edition), Nagano. Camera-ready in preparation.
Progress
1,537 specialist modules in the servable pool. Architecture validated at scale: adaptive routing correctly selects domain specialists for every input.
Traditional AI systems activate every parameter for every input — cost scales with model size. Modulith activates K specialist modules per position from a pool of N. Inference cost is constant regardless of how large the pool grows. Knowledge scales independently of compute. Validated: 1,537 modules in the servable pool, same inference cost as 4.
Independent training, frozen composition, automatic routing, cross-domain improvement, structural isolation, and capability control demonstrated across multiple scales.
1,537 specialist modules in the servable pool — code, mathematics, systems, chat, instructions, knowledge, and more. Trained under 5-layer quality gates and uniqueness pressure.
The Modulith Lab assessment pipeline is the first deployment of the composition architecture. Three independently trained frozen specialist modules compose via the same routing mechanism described above.
The Lean 4 formal specifications serve dual purpose:
• Define safety properties for the research programme
• Generate test prompts for the assessment engine
Fresh prompts auto-generated weekly from the same specs. The pipeline runs on consumer hardware (1.4GB VRAM) at zero marginal cost. The architecture's first proof: evaluating AI safety using modular composition.
On a three-step arithmetic problem, a composed team of four modules produced a reasoning step that none of its members produces alone. Each of the six modules involved was tested individually: none writes the step. The best single module reaches 2 of 6 points and then diverges. Removing any one seat from the composed team loses the step, including seats where the lead module is unchanged.
The routing policy selected the team from reward alone, with no information about which modules appear in known solutions. Its use of the relevant modules rose from chance to 86% over 640 attempts.
Bounds. One solve in 640 attempts. One problem. The answer is arithmetically correct and contains one step that is true but not meaningful — the scoring is mechanical and checks arithmetic, not reasoning quality.
Position paper accepted and presented at IEEE AIxSoftware 2026 (Summer Edition), Nagano. Formal verification programme for six safety properties. Machine-verified proof certificates in Lean 4.
Encrypted inference containers for regulated industries. Full attribution, guaranteed forgetting, EU AI Act compliance by construction.
Modulith Ltd — commercial entity. Patent holder. Licensing and deployment.
Modulith Research CIC — community interest company. Open research into modular AI architectures with provable safety properties. Issues all assessment reports independently.
Modulith Lab — production application of the formal verification framework. Independent AI model assessments against 24 Lean 4 verified properties. 18 models assessed. Reports issued by Modulith Research CIC.