
MLA400 PCIe Card packs four ARIES NPUs onto one PCIe Gen5 × 16 card, delivering 320 TOPS of AI inference at 135 W.
Install it into a standard rack server and run multi-stream video analytics, multi-user LLM inference, and multimodal workloads.

Specifications
NPU
Quad-ARIES (4× ARIES)
Performance
320 TOPS
NPU Frequency
1.25 GHz
NPU Cores
32
Memory Type
LPDDR4X
Memory Capacity
64 GB
Memory Bandwidth
Up to 266.8 GB/s
Flash Memory
32 MB
Host Interface
PCIe Gen5 × 16
TDP
135 W
Input Voltage
DC +12V (Slot + AUX 8-pin)
Dimensions
268 × 112 × 19 mm
Supported OS
Linux, Windows
Supported Frameworks
PyTorch, TensorFlow, TFLite, ONNX, Keras
SDK
SDK qb
UART
1 port
LED Indicators
Power × 3, Status × 4
Thermal profile options

Cooling Design
Heatsink + Active fan cooling
Target Deployment
Sustained throughput in ventilated environments
Dimensions
268 × 112 × 19 mm
Weight
834 g
Four ARIES,
one unified pipeline
Each MLA400 PCIe Card carries four ARIES accelerators, each with 8 cores running at 1.25 GHz. The 32 cores operate as a flexible compute pool with qb Runtime handling workload distribution, memory management and multi-chip scheduling across all four chips.
A single MLA400 PCIe Card can process 90+ concurrent video streams for vision workloads, or serve multiple simultaneous LLM users for generative AI deployment.

320 TOPS at 135 W
32 NPU cores deliver inference-dedicated throughput within a 135 W envelope. Models built and validated on GPU run on MLA400 PCIe Card through SDK qb.
Drop-In Server Deployment
Standard PCIe Gen5 × 16 interface and slim single slot form factor. Installs into off-the-shelf server platforms with no custom hardware.
Two Thermal Profiles
Active cooling for maximum sustained throughput, or Passive Slim for fanless environments where noise and airflow are constrained.

Deploy your model on MLA400 PCIe Card
MLA400 PCIe Card is the deployment platform for models that have already proven themselves in development. Run models at inference-optimized power efficiency in a production server environment.
qb Compiler converts GPU-validated models into native binaries, supporting PyTorch, TensorFlow, TFLite, ONNX and Keras.
Validated with 490+ models, pre-compiled open source models are available in Mobilint Model Zoo across vision and transformer to multimodal architectures.
Where MLA400 PCIe Fits
On-Premises LLM Serving
Serve large language models locally for enterprise teams. Data stays on-site with no per-query cloud cost.
Multi-Stream Video Analytics
Process dozens of concurrent camera feeds for facility-wide security, safety monitoring and quality inspection.
GPU-to-Production Deployment
Deploy models validated on GPU through SDK qb. The same model, now running at inference-optimized efficiency.
Edge Data Center
Add dedicated AI inference capacity to existing server infrastructure within a fraction of typical GPU power draw.


