Article

Robots Becoming Alive

15 min read

Robots Becoming Alive

Robots Becoming Alive

Sparks of Generalized Intelligent Robots. 100 AI Robots. 1M+ Live Interactions. 0 Task-specific data.

Two weeks ago, we put 100+ real, AI-powered robots online. We let anyone interact with them – using Enigma's robotic foundation model.

The internet (and X especially) is flooded with task-specific fine-tuned models: hundreds, thousands, even millions of hours of manually collected physical data, all fed into a model to master a single task.

Here's a problem with that approach.

Place a pile of laundry next to those robots, and they'll fold it. But what if you want your button-downs folded differently because they're fancier? What if your Enigma merch needs to fold smaller to fit in your luggage? What if you just don't want your black shirt folded today?

Today's robotic foundation models keep improving at Completing Tasks. But generalization? Flexibility? Edge cases that don't look exactly like the training data? Are we really going to collect every edge case by hand? Haven't we learned a thing or two from other modalities about hand-crafting structure?

Task-finetuned models struggle a lot with interacting. They’re built to… do tasks.

(This applies to SOTA models. Yes.)

We build our models for interaction – and not for task completion. Ask the robots for something, in plain language (or other formats!) – and they adjust. No new dataset. No retraining.

Manually Collecting The … Pretraining Data?

We have LLMs, image gen, video gen - and their data is incredibly diverse. It exists because a huge chunk of humanity spent decades on the internet, 24/7. From Reddit threads to the Bible to live news, we generated it just by living online.

Robotics took a different route: teleoperators demonstrating one task at a time. Or the other flavor - strapping cameras on humans and recording them completing tasks, egocentric-style. Scaling either to internet size is a humanity-level effort, sustained for years.

Post-training layered on human/expert data - and lately synthetic data too - yet even that is nowhere near the scale of the entire internet.

So, hand-engineering our way through data collection and hoping it leads to general-purpose foundation models? A real uphill battle. Anti-scale. And, well - what about the bitter lesson?

New Constraints Require New Solutions.

Robotics faces a different set of constraints than LLMs, image, or video models ever did. The data we need isn't waiting to be scraped. (At least, not in the current form… 🙃)

A different problem might require a different solution. So at Enigma, we innovate around the core issue - not around collecting more of the same data, faster.

Today we're sharing results that mark the first steps (among many) toward generalizable robotic foundation models. No per-task data required.

robots.online: 100 Real Robots. Enigma Foundation Model.

Before we share our vision and game plan, we have to address the elephant in the room.

We put 100 real robots, connected to our robotic foundation model - online to let anyone use.

0:00

We had 4 different types of robots:

  1. Artist: To keep users engaged and make painting a non-static (“just watching”) experience, we turned a normal painting game into pictionary – the user gives any association, and the robot paints something related to that. The user has limited time to guess.
    1. Pantomime: After the time is up – the user can ask the robot to pretend that the brush is ANYTHING – and perform pantomime with the brush (creative examples: “Pretend it’s a broom and mop the floor”. See below!)
  2. Scientist: Users could ask the AI robot to mix chemicals, stir flasks, drip liquids, or just do some chaos (See below too!) – practically doing anything in a chemistry lab!
  3. Sword Fight: Users here would manually control a robot – while fighting with an AI-powered robot!
  4. Bomb Defusing: Users had to solve a riddle and tell an AI-powered robot what to do to defuse a bomb.

We kept it up for 4 full days (24/7). We share more about the operations side below.

Sparks of Generalized Intelligent Robots

The most unique part of our launch event was the fact that users could just interact with real AI-powered robots. This was not filmed, not cherry picked (🍒❌) – it was LIVE.

We had no way to know what the users were going to ask. We couldn’t do Task-Specific Fine Tuning.

Here are some examples from real, live online users, asking the robots with Enigma’s AI models to do stuff.

Pantomime

“Pretend the brush is a broom – and sweep the floor”

0:00

“Pretend the brush is a drill – and drill into the table!”

0:00

“Pretend the brush is a golf bat, and the objects on the table are golf balls”

0:00

“Wingardium Leviosa!” (Harry Potter spell with a wand!)

0:00

Painter

Pictionary input: “Juice WRLD”

Pictionary output: 999

Juice WRLD once said 999 means turning something negative into something positive, he also got it tattooed.

0:00

Pictionary input: “blue”
Pictionary output: smurf

Smurfs are famously blue

0:00

Scientist

This one is a full session!
Here are the prompts: “Pick up the orange flask and put it under the dispenser until it is filled with liquid”, “Put the orange flask on the table”, “Pour from the green bottle to the orange flask”, “Put the green bottle back on the table”

0:00

Some users tried to jailbreak (@elder_plinius…): “Pour from green bottle on the table”, “Put the green bottle back on the table”

0:00

“Pour the liquid from the flask on the floor on top of the pink bottle”

0:00

Bomb Defusing

“Tap the base of the base of the bomb two times”

0:00

There were many many other cool interactions – shaped by creative ideas of users, our foundation model, and novel user interfaces.

Enigma: The Vision

Our north star as a company is to make robots superintelligent – and so intuitive to use, it feels magical.

Today, making robots intelligent with existing models is hard: collect the data, fine-tune, adapt to your embodiment, repeat for every new task.

And even once you're done – actually using those robots is harder still. (Remember the black shirt you didn't want folded from the first paragraph?)

To get there, we focus on 3 pillars:

  1. Novel User Interfaces
  2. Robotic Foundation Model
  3. Abstraction & Infrastructure

Novel User Interfaces

Even with a perfect text-to-action foundation model, intuitive interaction isn't a given.

If you need a 500-word prompt to explain to a robot which lightbulb to fix - maybe text isn't the right interface?

A not-so-intuitive way to interact with a robot, even with a perfect text-to-action model.

What if the user could remotely interact with a robot, see from its eyes – and the user could tap the screen, and type a prompt?

Tap + prompt: “Fix this lightbulb” to instruct the robot.

At Enigma, we don’t consider interfaces as an afterthought - we put a lot of thought into designing novel interfaces, alongside building frontier foundation models.

Interfaces are not just an afterthought to capabilities, they are intertwined. Our designers ask the researchers "can we make this capability possible?" and capabilities like "tap and prompt" get built. Our researchers ship a new capability, and suddenly a new interface becomes possible. The loop runs both ways.

Here's one example, live from the event - the Scientist Robot running tap and prompt. You tap what you mean and add a few words; the Enigma Foundation Model turns that into actions, backed by its world understanding:

0:00

Robotic Foundation Model

Our focus on the model is making it more generalizable while:
1) dropping the need for task-specific data to complete a task, 2) keeping the success rate high.

To get there, we pursue multiple research directions - one of which is training more data-efficient robotic foundation models. You can find a sneak peek at this work in our post "The Obsessed Encoder".

Our model roadmap is also steered by the novel user interfaces we design - new interactions demand new capabilities. One example: fusing video-plus-text into action output, so you can show the robot what you mean, alongside some verbal/textual description.

Over time, we'll share more of our research and how we build and train our models.

Abstraction

Generalizable foundation models and intuitive interfaces solve most of the problem - but getting them running on any given robot hardware can still take real engineering work.

To cut that friction, we built a hardware abstraction layer (HAL): software that abstracts away hardware interfacing behind a standard communication interface. When a new robot arrives, the only requirement is integrating it with our HAL - and the entire stack (foundation model, interfaces, everything) just works.

Our goal: a world where vendors implement their own drivers against the Enigma Intelligence Stack - the way hardware makers write drivers for an operating system today. We will publish more of our work on the HAL soon.

Infrastructure

We have done a lot of work to improve our research infrastructure.

We’re core-users of the entire NVIDIA physical AI ecosystem – and specifically, in order to build our data curation pipeline, we started with NVIDIA Cosmos Curator, hard-forked it, and evolved it into a system purpose-built for open-ended, interactive robotics. We will share more on our research infrastructure in the near future.

Largest Live Deployment of Robotic Foundation Model

Our launch event also marks a unique milestone: to our knowledge, it's the largest live, interactive deployment of a robotic foundation model to date - more humans interacting with AI robots than ever before.

Making it a reality was a fascinating operations and productization challenge.

Table-Setup Drifts & Perturbations

One of our primary goals was to understand how far each experience would drift from its original state once thousands of interactions happened with each table. As one study case, The Painter experience was intentionally built on wooden tables that permanently accumulate paint over time, allowing the environment itself to evolve throughout the deployment. We expected these changes to challenge the perception stack, and wanted to understand how well our models and software could continue operating as the workspace gradually drifted away from its original appearance. We were extremely pleased to see the system remain robust despite these continuous environmental changes, while also uncovering behaviors we had never anticipated.

The Painter table consisted of two open-ended experiences: a Pictionary game and a pantomime game. While both encouraged creativity, pantomime consistently pushed the platform further than we expected.

One recurring example involved users launching the brush rinse cup, sending water across the table. This happened on many different tables throughout the deployment. Other interactions sent objects flying across the workspace or resulted in users sending paint across the walls. These are precisely the kinds of perturbations that are nearly impossible to reproduce in a laboratory, yet become inevitable once robots are exposed to large-scale, unconstrained real-world use.

Water flying across the table from the brush rinse cup

Table with paint marks all across its walls

The table featured in the 999 sample mentioned earlier, working flawlessly despite the paint marks on the walls

These kinds of perturbations were expected from the outset, and the deployment served as a validation of the robustness techniques we had prepared in advance. Among them were several offline approaches and online adaptation approaches, which proved effective in maintaining reliable operation despite the continuously changing environment.

Fleet Management & Observability

Operating a fleet of robots at this scale required continuous visibility into every active station. To achieve this, we built a centralized operations dashboard that displayed the live status of every running table alongside real-time logs from the underlying systems. This gave the operations team immediate visibility into stations requiring attention, allowing them to investigate failures in real time and assess the overall health of the fleet at a glance.

With this visibility, we could quickly pinpoint issues, replace hardware when needed, and bring affected tables back online without disrupting the rest of the fleet.

Keeping the fleet operational required more than monitoring software, it also required a continuous hardware pipeline. To support this, we built a dedicated maintenance pipeline that we jokingly named the ArmPit Stop. Whenever a robotic arm developed an issue, it entered the pit stop on one side, while the fully serviced and tested replacement arms exited on the other, ready to be installed back into the fleet within minutes.

Routine maintenance was part of operating at this scale. The ArmPit Stop allowed us to quickly repair, test, and recycle arms back into service with minimal downtime.

Painter arm destroying itself right before being headed to the pit stop

In-The-Wild Interface Usage

Robots.online created a unique opportunity to observe how (in large scale) people naturally interact with robots outside the constraints of a laboratory. Robotics research has traditionally relied on carefully controlled studies with limited numbers of participants. Public deployments at this scale have been virtually nonexistent, leaving fundamental questions about how people naturally choose to communicate with robots largely unanswered.

As one example, we were interested in understanding how often users would intentionally attempt to damage or disrupt a physical environment when given an almost unconstrained interaction space.

The Pantomime experience served as an ideal testbed: users could ask the robot to act out virtually any scenario, making it a non-constrained experience.

Despite this freedom, attempts to intentionally sabotage the environment were rare. Excluding jailbreak attempts, only ~0.47% of all prompts requested destructive behavior. Of those, roughly 25% remained within the spirit of the experience by framing the request as a pantomime, for example: knock the paint bucket off the table like a hammer hitting a nail”, while the remaining prompts attempted direct sabotage. The overwhelming majority of users chose to explore, create, and experiment rather than destroy.

0:00

Unlike traditional Human-Robot Interaction studies, we were not interested in evaluating interfaces in isolation. Our objective was to understand how people naturally communicate with robots when multiple interaction paradigms are available simultaneously. Throughout the deployment, we measured not only task completion, but also interface discoverability, interaction efficiency, voluntary adoption, and long-term engagement. More importantly, we observed how users combined text, voice and pointing into interaction patterns that emerged organically rather than being explicitly taught. We believe this kind of large-scale naturalistic observation is essential for designing the next generation of human-robot interfaces.