We're building the intelligence layer for video
JARD builds compact video-language models that make continuous, long-form video understanding economically practical — the semantic layer between raw footage and AI agents.
Adam Gera, Co-Founder · JARD
The Bet
We started with an observation that is now almost consensus, and a conclusion that still isn't.
The observation: the world does not happen in text. It happens in motion. Language is how humans compress reality after the fact — powerful, portable, and lossy. A sentence can describe a collision, a defect on a line, a lecture, a surgery, an incident in aisle nine. But before the description there is the thing itself: shape, motion, sound, sequence. Signals that change over time. Causality lives in sequence, and video is sequence.
Most AI systems are trained on the compression. Video is the signal. Fine. Almost everyone agrees now.
Here is the part that hasn't been answered:
If video is the richest record of reality we have, why is virtually none of it economically practical to understand?
That gap is the entire company.
The world is not short on video intelligence demos. It is short on video intelligence anyone can afford to leave running. There is a difference between a model that can watch a clip and a system that can watch ten thousand streams, all day, every day, inside a real budget, at a real latency, under real privacy constraints. The first is a research result. The second is infrastructure.
Our bet is that video understanding becomes ubiquitous only when it becomes cheap enough to run continuously — and that getting there requires a different kind of model, not a bigger one.
So we are not trying to build the most intelligent model that has ever looked at a video. We are trying to build the most intelligence per watt, per dollar, per hour of footage — enough understanding to solve the problem, at a cost that lets you point it at everything.
That is the bet JARD is making.
What We Built
We built the company around three technical convictions.
First, perception that survives compression.
Long video is not an image problem repeated ten thousand times. Treating an hour of footage as one enormous prompt is the most expensive possible way to learn that nothing happened for fifty-eight minutes. A useful model has to do something harder: identify what changes, carry the relevant state forward, discard redundancy aggressively, and still retrieve the right history when a question arrives an hour later.
Our research targets exactly that temporal and computational bottleneck. Not a general claim about context windows — a specialized approach to representing and retrieving long-form temporal information efficiently, which is the specific thing standing between video and every system that wants to use it.
Second, memory instead of storage.
A model that only looks at video at query time doesn't know the archive. It is performing a temporary inspection, paying full price, every time, forever.
We built the opposite. Footage is understood once when it enters the system, converted into a durable semantic representation, and kept addressable at the exact second of the exact file. The archive stops being passive storage and becomes machine-readable memory — persistent across streams, devices, and locations, and available to agents that need to know what happened last Tuesday without re-watching last Tuesday.
Third, deployment as a first-class design constraint.
Most video does not want to leave the building. Bandwidth, latency, privacy, regulation, and plain cost all push perception toward the source. So efficiency is not a nice-to-have we optimize at the end — it determines the architecture. A compact model that runs in the cloud, in a private environment, or on edge hardware is a fundamentally different product from a large one that runs in exactly one place.
Perception, memory, and efficiency form a loop. We call the result a video intelligence layer, and the point of it is simple: make video computable.
We are not trying to replace frontier models, vector databases, or your application. We sit upstream of all of them. JARD interprets the video continuously and produces structured output; your database stores it, your agent reasons over it, and when a moment genuinely needs maximum general intelligence, you escalate that one clip — not the other 23 hours — to a frontier model.
Process everything efficiently. Escalate only the difficult moments.
That is the honest architecture. Treating every frame of every stream as a frontier-model request was never going to scale, and everyone building on video has already discovered this the expensive way.
Why Now
The last decade made text programmable. Language models turned words into tokens, tokens became the semantic layer, and everything downstream followed: documents became context, chats became workflows, code became executable knowledge.
Video has not had that moment.
The world's footage is still mostly dark matter to machines. It sits in archives, camera systems, factories, warehouses, classrooms, stadiums, clinics, drones, and drives — an enormous record of what actually happened, accessed almost entirely through filenames, folders, timestamps, and human memory. The richest record of reality is still outside the semantic layer that modern AI runs on.
Four things changed at once, and they changed recently.
Video stopped being media and became sensory input — for robots, industrial systems, security operations, inspection workflows, and education platforms. Multimodal and agentic applications moved from demos to roadmaps, and those agents need eyes. Users stopped accepting a fixed dashboard of detections and started expecting semantic answers to questions nobody pre-configured. And model compression, efficient architectures, and capable edge hardware finally made local semantic intelligence viable rather than theoretical.
So the demand exists and the silicon exists. What is missing is the layer in between.
The application layer is ready. The efficient video-understanding layer is not there yet.
There is an objection worth answering directly, because every serious person asks it: won't frontier inference simply get cheap enough?
Costs will keep falling. But the number of cameras, robots, agents, and continuously recorded workflows is growing at the same time, and the relevant question was never whether one video gets cheaper to process. It is whether a company can economically process thousands of continuous streams, retain long-term context across them, and still meet its latency and privacy requirements. Specialized infrastructure stays valuable in that world — arguably it becomes more valuable, because the volume it has to absorb keeps growing.
Who This Is For
We are starting with AI-native companies whose products already depend on video.
These teams have strong engineers. What they do not have is a reason to spend two years building and maintaining their own long-form video model stack, when their actual differentiation is the workflow, the customer experience, and the industry data on top.
The pain is sharpest right now in industrial and warehouse intelligence, robotics and physical AI, retail and security analytics, inspection and incident analysis, and recorded knowledge work like lectures and training. What these have in common is not an industry — it's a shape: video that never stops, budgets that do, and questions that keep changing.
It is worth being equally clear about what we are not building. Not a video-generation product. Not a vector database. Not a labeling platform. Not a generic CCTV dashboard. Not a different bespoke application for every industry that asks.
We build the model, the memory, the API, and the runtime. Everything above that, our customers build better than we would.
Where We Are
We are early, and we would rather say so plainly than dress up interest as traction.
What exists today is JARD Spark, our compact video-language model, in private technical evaluation. We are working with five design partners across education, warehouse intelligence, retail analytics, CCTV, and deep note annotation — not labeling — at stages ranging from technical evaluation to scoped pilots, on production workloads where the current economics are visibly broken. We are not currently raising.
The next twelve months are designed around one testable claim rather than a broad narrative: that we can take a real continuous-video workload, meet the quality threshold that application actually requires, and do it at materially better economics — then repeat it across several customers.
Everything we measure ladders up to that. Verified quality on long-form tasks customers care about. Cost per processed video hour against an agreed baseline. Stable throughput and latency on target hardware. Reliability over long, continuous streams. And the one number that combines adoption with economics: paid production video hours processed through JARD, subject to quality and margin thresholds.
If that number compounds, the thesis is working. If it doesn't, we would rather find out early and loudly.
Where This Goes
In five years, JARD should not be describable as "a small video model."
It should be the semantic operating layer for the visual world — the infrastructure that developers and agents use to turn raw footage into events, state, memory, and context that software can search, reason over, and act on.
The path is deliberate: a compact long-form video model, then developer APIs and an edge runtime, then persistent memory across streams and devices and locations, then the standard video context layer for agents and physical AI.
As agents move out of browsers and documents and into factories, stores, warehouses, classrooms, robots, and wearables, they will need continuous visual context to be useful at all. Someone has to supply it. We would like that to be a layer everyone can build on rather than a capability only the largest labs can afford to run.
Language models made text actionable. JARD makes video actionable.
Come Build With Us
We are looking for two kinds of people.
Design partners.
Teams running real, continuous video workloads who are willing to put us against their current pipeline on their own data, with a clear success metric and an honest comparison. If your inference bill scales with footage and that is starting to hurt, we should talk about the workload.
Builders.
Researchers and engineers working on efficient architectures, temporal modeling, inference optimization, and edge runtimes — plus the product and commercial people who turn that into something teams can actually deploy. If you believe the next frontier of AI won't be limited to what humans have written down, but built from what actually happened, we are hiring.
The visual world is the largest untapped input to AI. Making it usable is the work.
— Adam Gera, Co-Founder, JARD