โ† All work WIP ยท being written

Case Study ยท Hoomanlabs ยท Voice AI Platform

Making the conversation visible

A full redesign of the four surfaces that ops leads, collections managers, and support heads use to run AI voice agents across millions of calls, plus the migration out of Mantine into shadcn/ui. Sole designer, owned end to end from research to launch. The goal: let someone who has never written a prompt run a working agent, on a product that finally looked like one product.

Role
Product Designer โ€” sole designer, research through launch
Company
Hoomanlabs
Surface
Agent Builder ยท Call Logs ยท Campaigns ยท Insights ยท Design system
Status
Live in production, still shipping
Hero โ€” the redesigned Agent Builder (drop screenshot later)
The redesigned Agent Builder. A collections agent, greeting, identity check, promise-to-pay. The expanded node shows what the agent says and what it listens for, written the way you'd brief a new hire. No prompt box, no JSON, no engineer.

Context

Hoomanlabs sells AI voice agents that make and take phone calls at scale: debt collection, lead qualification, support, recruitment screening. The platform has run [10M+] calls across lending, e-commerce, travel, and logistics.

The people on these pages are not who the category was built for. They're collections leads at NBFCs, support managers at D2C brands, recruiters screening at volume โ€” operators with a number to hit and no engineering budget to hit it with. They'd been sold on replacing a call centre. Then they opened a product that assumed they could write a system prompt.

The product already existed and already had customers. That's the harder brief: four live surfaces to redesign without breaking the people already depending on them โ€” and a component library underneath that had quietly become the ceiling on how good any of them could get.

The problem

Three problems, and the third is the reason the first two kept not getting fixed.

The builder asked people to think in the wrong shape. A call is a graph, but operators don't think in graphs โ€” they think in scripts. First I greet them, then I check I've got the right person, then it depends. The product asked them to declare that structure up front, in the abstract, before hearing a single call. For a developer that's fine. For a collections lead it's a blank page with stakes.

A bad call and its cause lived on different planets. Call logs told you what happened on one call. Usage dashboards told you aggregate minutes. Neither answered the only question an ops manager actually has โ€” which part of my script is losing people, and what do I change โ€” and nothing connected a call that went wrong to the place you'd go to fix it.

And the four surfaces didn't look related. Mantine was the right call when the team picked it; it bought speed at the stage where speed was the whole strategy. But every screen carried its defaults, and years of one-off overrides on top had drifted the surfaces apart. Design work kept terminating at the same wall โ€” a spec came back as "the library can't do that," a compromise shipped, and over enough compromises the compromises were the product. The data-dense screens suffered worst: call logs and campaigns needed density and custom cells the library's table wouldn't give.

Operators don't think in graphs. They think in scripts, and the product was asking them to think in the wrong shape.

The insight that drove the redesign
[One real sentence from a customer or ops lead. The single highest-value line you can add to this page.]
โ€” [Ops lead, lending customer, paraphrased]

Why now

Every new customer needed hand-holding through their first agent, which doesn't scale past a handful of logos. The voice models had gotten good enough that the interface was the bottleneck, not the AI. And the redesign was the only moment the migration would ever get funded โ€” nobody approves a quarter of "rewrite the components and change nothing." Coupled to a redesign, it's just how the redesign ships.

Research & framing

I interviewed ops leads at customer accounts directly, and worked backwards through support tickets and what the founders and CS team already knew โ€” the recurring questions nobody had connected into a pattern yet.

I also studied how Bland, Vapi, Retell, ElevenLabs Agents, and [regional players โ€” name them] structure agent creation. Nearly all of them design for a developer building on behalf of a customer. Hoomanlabs' users had no developer. That gap was the opportunity.

Then I walked the full lifecycle: an operator writes an agent, attaches a contact list, launches a campaign, reviews the calls that went badly, edits, relaunches. The reframe that reset everything: the loop is the product, not the launch. Creation was the moment the product had been designed for. Repair is the moment users actually live in.

2026|Reframing, who's actually here?

Old framing

  • An agent is a prompt to be configured
  • Show the model's capabilities; let the user assemble them
  • Language of prompts, tokens, LLM settings
  • Success is creating an agent
  • Reviewing calls and editing agents are separate jobs
  • The component library decides what design is possible

New framing

  • An agent is a script, written the way you'd brief a person
  • Show the conversation; hide the machinery behind it
  • Language of scripts, callers, outcomes, next actions
  • Success is a campaign that ran and got fixed
  • Reviewing a bad call is one click from fixing its cause
  • We own the components, so design sets the ceiling

Design principles

  • Show the conversation, not the configuration. If it can be said the way you'd say it to a person, don't express it as a setting.
  • A bad call points at its own cause. Review and repair belong on the same rail.
  • Density is a feature, not a compromise. These are people reading hundreds of calls a day, not browsing.
  • Own the primitives. A component we can't change is a design decision someone else already made for us.

User flow

The full agent lifecycle โ€” blank canvas to a campaign that's actually working, plus the recovery path that closes the loop when it isn't. The dark nodes are the additions that turned a one-way launch into a loop.

OperatorLive callIntroduced in redesign1 ยท Build agent2 ยท Attach contacts3 ยท Launch campaign4 ยท Live callsOutcomeConnected5 ยท Review calltranscript linked to node6 ยท Edit the failed nodefix what failed ยท relaunch

Hover or tap any node to see what it does.

Dark nodes (5, 6) are the additions that made launching a loop instead of a one-way trip. Hover or tap any node.

Design solution

Three acts: the Agent Builder, the operational surfaces around a live campaign, and the systemic move โ€” replacing the foundation all of it stands on.

Act 1 ยท Agent Builder

Making a conversation something you can see

The canvas now starts from the happy path, not an empty graph. A new agent opens with a spine already laid down โ€” greet, verify, ask, close โ€” and branches grow out of it as the operator remembers the ways a real call goes sideways.

Each node holds two things: what the agent says, and what it listens for. The instructions field is framed as briefing a new hire rather than prompting a model โ€” same underlying content, and it changed the sentences people wrote. [What changed, if you noticed it.]

The version that didn't work: a wizard. Before the canvas, I explored a step-by-step form builder โ€” answer a sequence of questions, get an agent. It tested well right up until a conversation branched, which is immediately. Real calls fork on the second turn: wrong person, no answer, call me later. A wizard can express a sequence and not a shape, and a phone call is a shape. Killing it is what made the canvas non-negotiable.

Agent Builder โ€” canvas + expanded node (drop later)
The wizard exploration, killed (drop later)
Act 2 ยท Call console, Campaigns & Insights

From one call to a thousand

Transcript, recording, summary, and outcome on a single screen โ€” the transcript as the spine, everything else anchored to timestamps in it.

The move that mattered: each transcript turn links back to the builder node that produced it. Reviewing a bad call is one click from fixing its cause. That link is what makes the loop in the flow diagram real rather than aspirational, and it's the improvement I'd point to first.

Hinglish shaped the transcript design. Calls code-switch constantly โ€” a sentence starts in Hindi and lands in English. I render Hindi in Devanagari and English in Latin, mixed inline, exactly as the caller spoke it. Romanizing everything would have made a tidier column and a less true one; an ops lead scanning a transcript is checking whether the agent sounded right, and a transliterated sentence doesn't let you hear it. [What you traded to get this โ€” font stack, line height, mixed-script alignment.]

Campaigns wrapped the launch itself: contact upload, scheduling, retries, pacing, live monitoring. [Compliance constraints โ€” calling windows, DND โ€” if they shaped the design.]

Insights answers the aggregate question the same way the call console answers the single one: drop-off by conversation stage, plotted against the flow the operator built rather than against abstract KPIs. A chart reading "62% of calls end at the identity-check node" names the problem, names the node, and puts the fix one click away. "Average handle time: 47s" names nothing. [Confirm this is the shape you shipped.]

Call console โ€” transcript linked to nodes (drop later)
Insights โ€” drop-off by conversation stage (drop later)
Act 3 ยท The foundation

Mantine โ†’ shadcn/ui, without a rewrite quarter

The systemic move. Mantine is a library you theme; shadcn is source you own โ€” Radix primitives underneath for accessible behaviour, Tailwind tokens on top, and component files sitting in the repo where they can be edited. The practical difference: the design system stopped being a Figma file nobody opened and became variables the product actually reads.

Three things forced it. The surfaces had drifted apart โ€” defaults plus years of overrides had left four screens that didn't look like one product. The theming ceiling meant Hoomanlabs couldn't look like Hoomanlabs. And the data-dense screens โ€” call logs, campaigns โ€” needed density and custom cells the library's table wouldn't give at any amount of configuration.

The strategy was to never let it become a separate project. One rule held it together:

No screen migrated without being redesigned. No redesign shipped on Mantine.

That coupling is why it's still moving. Migration work is the easiest thing in the world to defund halfway; attached to a redesign customers were waiting for, it has cover the whole way. We went surface by surface rather than big-bang, which kept every release shippable and meant the product was never half-broken in front of customers. [Which surface first, and why.]

The honest cost. We own the maintenance now. There's no upstream release that fixes our components while we sleep, and accessibility is our problem rather than a vendor's promise. [What you did about that.] Worth it, but it's a trade, not a free upgrade.

Process & iteration

Sole designer across all four surfaces, working with [N] engineers who built it. Owning research through launch meant the reframe survived the trip โ€” nothing got translated into someone else's summary between the customer interview and the shipped screen. It also meant every unmade decision was mine, and there were a lot of them.

[N] iteration rounds per surface. The wizard exploration above was the biggest thing I killed; [anything else worth naming].

Impact

No clean before/after metrics โ€” instrumentation landed after the surfaces did. What I have instead:

  • The surfaces shipped, redesigned and migrated, live in production and still rolling forward.
  • The product stopped looking borrowed. Four surfaces that read as one product instead of four eras of one.
  • Sales changed how they demo. [What specifically changed โ€” live builds instead of slides, a deal that turned on it.]
  • Onboarding stopped needing hand-holding for a customer's first agent. [Or the specific support question that disappeared.]
  • Design velocity changed shape. New surfaces stopped starting with "can the library do this."
  • [The real customer quote goes here.]

Reflection

If I ran it again I'd migrate the shared primitives before touching any surface. Going surface-first put the user pain first, which felt right and meant the earliest screens got built on components that were still moving โ€” some of them got redone. The boring layer first would have felt slower for a few weeks and been faster by the end.

The lesson that generalised: a redesign is the only affordable moment to replace a foundation, but inside that window the foundation still has to go first. Couple the two efforts to get the migration funded, then sequence them the other way round.

Role Sole designer. Led the redesign across all four surfaces and the Mantine โ†’ shadcn migration โ€” research, flows, screens, and the system underneath, through to launch.
Collaborators [Founders, engineering, CS.]
Stack Figma, shadcn/ui, Radix, Tailwind.