top of page
MLA400 PCIe Card_Active.jpg

MLA400 PCIe Card

Quad-ARIES AI inference accelerator for on-premises deployment

MLA400 PCIe Card packs four ARIES NPUs onto one PCIe Gen5 × 16 card, delivering 320 TOPS of AI inference at 135 W.


Install it into a standard rack server and run multi-stream video analytics, multi-user LLM inference, and multimodal workloads.

MLA400-PCIe-Card_Set_Server.jpg

Specifications

NPU

Quad-ARIES (4× ARIES)

Performance

320 TOPS

NPU Frequency

1.25 GHz

NPU Cores

32

Memory Type

LPDDR4X

Memory Capacity

64 GB

Memory Bandwidth

Up to 266.8 GB/s

Flash Memory

32 MB

Host Interface

PCIe Gen5 × 16

TDP

135 W

Input Voltage

DC +12V (Slot + AUX 8-pin)

Dimensions

268 × 112 × 19 mm

Supported OS

Linux, Windows

Supported Frameworks

PyTorch, TensorFlow, TFLite, ONNX, Keras

SDK

SDK qb

UART

1 port

LED Indicators

Power × 3, Status × 4

Thermal profile options

MLA400_01.png

Cooling Design

Heatsink + Active fan cooling

Target Deployment

Sustained throughput in ventilated environments

Dimensions

268 × 112 × 19 mm

Weight

834 g

Four ARIES,
one unified pipeline

Each MLA400 PCIe Card carries four ARIES accelerators, each with 8 cores running at 1.25 GHz. The 32 cores operate as a flexible compute pool with qb Runtime handling workload distribution, memory management and multi-chip scheduling across all four chips.

A single MLA400 PCIe Card can process 90+ concurrent video streams for vision workloads, or serve multiple simultaneous LLM users for generative AI deployment.

MLA400_01.png

320 TOPS at 135 W

32 NPU cores deliver inference-dedicated throughput within a 135 W envelope. Models built and validated on GPU run on MLA400 PCIe Card through SDK qb.

Drop-In Server Deployment

Standard PCIe Gen5 × 16 interface and slim single slot form factor. Installs into off-the-shelf server platforms with no custom hardware.

Two Thermal Profiles

​Active cooling for maximum sustained throughput, or Passive Slim for fanless environments where noise and airflow are constrained.

MLA400-PCIe-Card_Set.jpg

Deploy your model on MLA400 PCIe Card

MLA400 PCIe Card is the deployment platform for models that have already proven themselves in development. Run models at inference-optimized power efficiency in a production server environment.

qb Compiler converts GPU-validated models into native binaries, supporting PyTorch, TensorFlow, TFLite, ONNX and Keras.

Validated with 490+ models, pre-compiled open source models are available in Mobilint Model Zoo across vision and transformer to multimodal architectures.

Where MLA400 PCIe Fits

On-Premises LLM Serving

Serve large language models locally for enterprise teams. Data stays on-site with no per-query cloud cost.

Multi-Stream Video Analytics

Process dozens of concurrent camera feeds for facility-wide security, safety monitoring and quality inspection.

GPU-to-Production Deployment

Deploy models validated on GPU through SDK qb. The same model, now running at inference-optimized efficiency.

Edge Data Center

Add dedicated AI inference capacity to existing server infrastructure within a fraction of typical GPU power draw.

Stock Image.jpeg

AI-enabling 4,000 cameras without infrastructure replacement

Mobilint ✕ IntelliVIX

MLA400 PCIe Card-based AI servers power real-time inference across 4,000 video feeds from highway cameras, while Intellivix's Vision AI and VLM platform automatically detects and verifies highway incidents in Korea Expressway Corporation's AI CCTV transformation project.

Products_Resource

Contact our team today for more information on our products and solutions

bottom of page