NewChunking Qwen3.5's gated DeltaNet for 1.5x faster prefill on Apple Silicon

Models & infrastructure for
physical AI, robotics, autonomous systems, drones, warehouse automation, on-device intelligence

Run any model on any edge chip, with builds that beat the vendor toolchains, staged rollouts over the air, and telemetry from every device in the field.

The Operating System for Physical AI

Ship models to your fleet the way you ship software. Prysm compiles a build for every chip, benchmarks it on the real hardware, rolls it out over the air, and watches it run in the field.

Compile

Compile models into optimized builds for every chip in your fleet.

model.onnx → .prysmsha 7b41a9
hailo-8int8 · chip-nms84s
jetson-orinfp16 · trt112s
kria-k26int8 · dpu97s
m4-maxfp16 · metalqueued
Benchmark

Benchmark every build on the chip it ships to.

Throughput0.0fps
Latency0.00ms
Power0.00W
yolov11n · jetson orin nano · fp16
Deploy

Push model updates to every device in your fleet. Tag, version, roll back.

v2.3.0v2.3.1rolling out
canary2 / 2
warehouse-a12 / 12
warehouse-b8 / 14
roll back to v2.3.0
Monitor

Trace every inference in the field. Catch drift before it becomes a problem.

throughput fpslatency p99 ms
Compile

Deep compiler research for more throughput on less hardware

Most fleets ship the vendor's default build, because hand-tuning every model for every chip doesn't scale. Prysm's compiler does that work on every build.

Faster than the vendor's own build

Up to twice the throughput of the vendor's default build, measured with the same model on the same chip.

Throughput131.1fps
Latency e2e7.62ms
Power1.15W
yolov11n@416 · hailo-8 · chip-side nms

Run the same workload on less hardware

At 197 frames per second, one Jetson Orin Nano carries six 30 fps camera streams. Faster builds cut the device count for the same job.

jetson orin nano · 197.0 fps
cam-030 fps
cam-130 fps
cam-230 fps
cam-330 fps
cam-430 fps
cam-530 fps
six streams · one module

Deployments feed the compiler

Every deployment feeds performance data back into the compiler and fine-tuning data back into the models. Your fleet gets faster the longer it runs.

fleet
flywheel
CompileBenchmarkDeployMonitor
Deploy

Deploy over the air.
Run at full speed.

Model delivery for production fleets. Push updates reliably at fleet scale, across sites, chips, and versions.

ModelTaskInputParams
YOLOv11-ndetection416×4162.6 MDeploy ↗
ResNet-50classification224×22425.6 MDeploy ↗
MiDaS-smalldepth256×25621.3 MDeploy ↗
Whisper-tiny (encoder)speech10 s window7.8 MDeploy ↗
Runs on
Monitor

Watch the whole fleet from one screen.

Search any inference, replay any run, and catch problems while they are still small.

Inferences
41.2M
last 30 days
Success rate
99.97%
312 faults
Latency p99
7.9 ms
fleet-wide
Fleet power
38 W
28 devices
Inferences 41.2M
Success rate 99.97%
Latency 7.9 ms p99
p99p90p75p50
Accuracy proxy mAP 0.288

Trace every inference

Inputs, outputs, timings, and device state, tied to the exact build that produced them.

Fleet metrics per build

Track throughput, latency, power, and accuracy per device and per version.

Search and replay

Filter by device, model, version, or fault. Replay any run exactly.

Tie every trace to a build

See which build produced every number, on any device in the fleet.

Models

Edge-aware foundation models
built for the field

State-of-the-art task-agnostic models, co-developed with the compiler and runtime, so they run at full performance on the chips your fleet actually ships.

Learn more →
person 0.94
forklift 0.88
pallet 0.91

Designed for quantization

Designed for INT8 from the start. Accuracy holds through quantization and compilation.

Fine-tuned on your fleet's data

Telemetry from your devices feeds fine-tuning, and the model specializes on the scenes your cameras actually see.

Built with the compiler

Model and compiler are designed together, so the network uses everything the chip has.

Retrained as your fleet evolves

Add a site or swap cameras and the next training cycle picks it up.

Procure

Find your chip.

Set power, weight, price, and frame-rate constraints, and get the chip that clears them, measured on real hardware.

Set your constraints

Cap power, weight, and price, set a floor on FPS, and the matrix returns the targets that qualify and the one that wins.

Every model, every chip

Detection, depth, and speech models measured across every supported target.

Decide before you buy

Compare chips at the R&D stage, before you commit to hardware.

Contact

Ship AI to the edge.

Deploying models in the field? We should talk.

Open source

Alloy

GPU compute & LLM serving for Apple Silicon.