Skip to content
AI Research

AI Research

Running inference on constrained devices, and what that costs in memory, latency and power.

1 published

BenchmarkMay 2025AI Research

EAI 0.9: INT4 LLM Runtime — 11 tok/s on Cortex-M85

EAI's new quantized inference path squeezes a 1.3B-parameter model into 312 MB of flash and runs at interactive speed on a 480 MHz microcontroller. Block-streamed weight loading reduces peak SRAM by 94%.

EAILLMBenchmark
Read