

Joint AI Research & Dynamics
JARD
The intelligence layer for video. Continuous understanding, at a cost that lets you point it at everything.
What we are building
The world's footage is still dark matter to machines. There is no shortage of video intelligence demos. There is a shortage of video intelligence anyone can afford to leave running.
An hour of footage is not a prompt. A useful model has to identify what changes, carry the relevant state forward, discard redundancy, and still retrieve the right second when a question arrives an hour later.
JARD understands footage once as it enters the system and keeps it addressable at the exact second of the exact file. Routine perception runs locally; only the moments that genuinely need frontier reasoning get escalated.
Principles
- 01Intelligence per watt, per dollar, per hour of footage — not the biggest model.
- 02Understand footage once, then keep it as memory — not storage re-inspected at full price with every query.
- 03Deployment is a design constraint: cloud, private, or on-device is the same model.
What the model outputs
One model, steerable through a prompt. Ask for any of the below, in the shape and detail you need.
We sit upstream of your database, your agent, and your application — we produce the structured context, you build the product.
- Timestamps, scenes, segments
- Objects, people, tools, locations
- Actions, events, intent, state changes
- Summaries, QA pairs, retrieval indexes
Under 2% the size. Still ahead on video understanding.
JARD Spark is a compact video model, less than 2% the size of today's flagships, yet it leads on the standard video-understanding suites. Smaller models make it economically practical to leave video understanding running across thousands of continuous streams, not just process one clip more cheaply.
| Benchmark | JARD Spark | Gemini 3 Pro | GPT-5.2 | Claude 4.5 Opus | Nemotron 3 Nano Omni |
|---|---|---|---|---|---|
| VideoMME | 88.5 | 87.7 | 85.8 | 81.4 | 72.2 |
| MLVU | 87.1 | 83.0 | 85.6 | 81.7 | - |
| MVBench | 78.1 | 74.1 | 78.1 | 67.2 | - |
| Avg (accuracy) | 84.6 | 81.6 | 83.2 | 76.8 | - |
| F1 benchmark | |||||
| UCF Crime (F1) | 97.32 | 97.12 | 93.1 | 91.7 | - |
Scores reported on public eval splits. Accuracy benchmarks are averaged separately; UCF Crime is reported in F1 and is not included in the accuracy average. JARD Spark is a preview build; final numbers may vary.
AI that truly understands video.
A live look at the video intelligence layer — upload hours of footage, then ask anything.
Any length · Rich visual understanding · Interactive agent · Reports & insights
Built for learning, research, marketing, security and sports.
A glimpse of what the underlying model can do.

Let's talk
Running a real, continuous video workload? Building efficient architectures, temporal models, or edge runtimes? We'd like to hear from you.

