Fresh Labs
Jinny AI

Jinny-Hero-2 1

From rigid commands to natural conversation

Background

Human-robot interaction has traditionally been rigid and transactional. Fresh Labs set out to change the status quo. By pairing the language understanding of ChatGPT with the generative voice technology of ElevenLabs, we built Jinny: a proof of concept for a robot that listens, understands, and responds the way a person would.

Challenges

Getting a robot to hold a natural conversation raises problems that go beyond speech recognition:

  • Telling conversation from commands. A robot that treats every sentence as an instruction is exhausting to talk to. Jinny needed to tell the difference between “how was your day” and “move forward three feet.”
  • Sounding human, not robotic. Text-to-speech that sounds synthetic breaks the illusion immediately. The voice had to carry natural rhythm and inflection.
  • Building fast, on a lean team. With a four-week timeline and a small cross-functional group, the proof of concept had to prove real value without the runway of a full product build.

Key Contributors

Elisha Terada
Elisha Terada
Johnny Rodriguez
Johnny Rodriguez
Mikey Weller
Mikey Weller
Steve Yin
Steve Yin
Jinny-Architecture – 01
Jinny-Foreground – 01
Setting the direction

Brainstorming a robot that could actually talk

Several remote sessions set the direction for the project: not just a robot that could move, and not just software that could talk, but a robot that could do both at once, changing the nature of human-robot interaction.

The team broke the problem into three tracks—robot hardware, interface, and conversational software—each demanding its own tools and expertise, and each essential to the illusion of a robot that felt present in the conversation rather than just responding to it.

Ideation in parallel

Prototyping and development toward one shared goal

Once the team agreed on a simple prototype that could hold a natural conversation and perform basic physical tasks—moving around a table, responding to its environment—individual members began building their pieces independently. Progressing on separate tracks, continuous check-ins ensured every piece was aligned with the same integrated outcome.

  • Interface design: A lightweight interface running on a Google Pixel 6a paired simple on-screen text and expression with Jinny’s spoken responses, giving users a visual anchor to go along with natural conversation.
  • Robot hardware: The mBot2 Neo’s built-in mobility and sensors let Jinny move and react physically on a table, without the team needing to design and build a chassis from scratch.
  • Conversational software: A Function Calling API let Jinny distinguish between a passing remark and an actual instruction—a chat, not a command line—while Speechly and ChatGPT handled recognizing and understanding what was said.
Steve-and-Jinny-img 01
Putting Jinny on display

From lunchtime demo to industry launch

Jinny’s first audience was internal—a lunchtime showcase for Fresh colleagues. From there, the team took Jinny to the Northwest Robotics Alliance event focused on generative AI in robotics, where it drew a positive response from attendees working in the same space.

What started as a four-week proof of concept became a working demonstration of opportunities in human-robot collaboration that can be showcased via rapid prototyping, proving a concept at a smaller scale before expanding the effort.


Results

3

Generative AI technologies integrated

2

Monthly recurring subscribers acquired

4

Weeks to deliver proof-of-concept