AI research lab

Puma AI Private local inference, agent harnesses, and memory store.

We study how useful AI can run closer to the user — practical methods for private local models, observable agent workflows, and durable memory stores.

Local-first
private inference
Observable
agent harnesses
Persistent
memory store

Private inference

Finding model, runtime, and hardware combinations that keep sensitive context close to the user.

Agent harnesses

Building evaluation loops that make tool use, state transitions, decisions, and recovery inspectable.

Memory store

Building durable memory that agents can write, retrieve, and inspect across long-running work.

Long-term vision

AI that is private, steerable, and safe

Private compute

More intelligence should run locally, with cloud services used deliberately and visibly.

Inspectable workflows

Agent systems should expose decisions, failures, permissions, and tradeoffs to their users.

Memory store

Useful agents need memory that persists beyond a single prompt and stays under user control.

Backed by builders

Angel and pre-seed support from leaders in crypto and AI.

Work with us

Help make local AI practical

We are interested in engineers and researchers who care about private inference, reliable agent systems, and unusually good test environments.