Expert ArticlesAI Applications & Use Cases

Thoughts on Voice AI: Why Pricing for Value May Not Be Enough

By Ash Kate
Thoughts on Voice AI: Why Pricing for Value May Not Be Enough

Article content

Thoughts on Voice AI - pricing, business models, survival. Let's price for Value and not per minute, shall we?

By Gregory Ipe, Head of Business, Hunar.AI

LinkedIn is currently debating whether Voice AI should cost ₹2 per minute, ₹1.50—or eventually even less. At an infrastructure level, the direction of pricing is probably right. Speech recognition, LLM inference, text-to-speech and telephony will keep getting cheaper.

A lot of the same posts also argue that price per minute is the wrong way to look at it. We must price for VALUE.

That is, somehow start earning a proportion of the underlying economic value unlocked.

This sounds logical on the surface and why not - pick a sector that has a high SVP (sectoral value per employee), provide increased productivity per employee* and look to take home a portion of the savings.

But this misses the practicality point, and there are a few structural reasons applying downward pressure on Voice AI pricing despite the value unlocked :


1) Price is becoming the differentiator 

Most Voice AI providers seem to own the orchestration layer with a combination of modular providers - STT, Noise Filters, Inference, TTS, telephony. Everyone seems to optimise for latency, interruptions, accents and other nuances.

Most results of pilots seem to be in the ballpark of the same set of results metrics ie the differentiation between one vendor to the other is not obvious at first. The only differentiation would be ROI and hence downward pressure on price.

*Voice AI is fundamentally a productivity layer on human engagement effort. The higher the value that arises out of that effort the higher the SVP. Think of a sophisticated tele-caller that AmEx employs versus a standard BPO employee.

And the prize to win is large. - Gregory Ipe


2) Procurement as a function exists to put downward pressure

In an early maturity market, enterprises with AI Implementation teams and defined charters in the AOP would be the first buyers. And procurement functions exist here solely with an objective to compare rates and exert downward pressure.

No ROI discussion - because ROI is specific to the enterprises’ workflow and prices are decided before the pilot even starts.

In a mature market even SMBs are broadly aware of the average rates in the market and hence the SMB as a route to profitability is a myth. SMBs do not automatically solve the problem. A smaller contract with enterprise-style implementation is worse economics. SMBs become attractive only when the product removes most of that implementation effort.

Enterprises seem to be the way forward (for scale, and this is nuanced, bear with me for now) and this creates downward pressure on price.

The logic would be to invest now into enterprises - lower price - increase minutes - obtain better supplier negotiated rates - lower price price - scale further and so on. This bets on staying power and once the dust settles assume that scale is the moat.


3) FDE / Support enabled GTM causes pressure to “scale”

Typically deployments at enterprises are outcome oriented with pilots failing because “we didn’t get the conversions that we expected”.

What they’re actually saying is that somehow Voice AI companies need to figure out not just Agent workflow design, prompt iterations but also downstream HITL workflows and the related change management. This tends to be effort intensive. Let’s look at the math at play.

Take the case of a company that operates at INR 2 per minute with gross margins of 10% (extreme). To service enterprises - with assuming some dedicated account management and a dedicated FDE - let’s say the CTC would be about 2L cash outflow a month.

To just break even at a unit accounting level you’d need a billing of INR 20 per month. Ie a million minutes until this unit breaks even. And this is even before you fund for sales, engineering, R&D, Marketing and other overheads. Brutal.

To breakeven you'd need a million minutes per FDE. How many companies manage that?

This creates the need to scale quickly and find low complexity - high volume use-cases, which is going to be the jugular every Voice AI company goes for. This creates additional incentives to drop prices to get in the door.


Okay so what needs to be done to escape this ?

There are no clear answers here.

But there seem to be prudent decisions that Voice AI companies will take to survive. I attempt to provide broad directional markers here and my views on how the stack evolves. Voice AI actually morphs into 4 business each with it’s own pricing models.

And how does one understand what is the best model that’s suited for the company?

The imperative first would be to identify the actual role that Voice AI plays in the sector of choice.

And a useful heuristic is to identify Modes of operation and Mission statements for the voice AI.


Modes

This is HOW the voice AI is deployed within the Org / Who owns the Execution? There are 2 broad modes in which Voice AI can operate :

1) Standalone Mode - Voice AI exists in a role that a human cannot handle today. These are typically lower involvement mission (below) statements where AI can do the job to about 70% of what a human would’ve done.

Use cases here would be 1) After hours support to handle incoming / CX 2) Nudge use cases or 3) Information dumping or Appointment confirmation use-cases.

2) Co-Pilot or HITL (Human in the Loop) Mode - Voice AI exists to make the human employee’s productivity better.

Typical LQ (Lead Qualification) or Screening use-cases. There is a human that is downstream to a set of leads that the AI Agent qualifies. This model involves technically the “double-cost” issue of both the human and the agent and this would work when you significantly reduce the count of the Humans in the loop.

Such use-cases would always run into Before AI - After AI comparisons at a cost per hire / cost per onboarding / cost per acquisition level. The pressure on cost is going to be brutal here.


Mission Statements

This is WHY the customer wants to deploy Voice AI / What economic value is desired.

Typically there are:

1) Efficiency improvement or Reduce Cost missions - These are mission statements which are looking for an arbitrage on human labour. Cost reduction OR TAT reduction for a fixed scope of work, in short.

Take for eg. I have 25 HR resources employed only to call candidates up and check whether they will turn up for the walk-in job drive. These are a fixed number of openings I need to hire for and the more number of candidates who show up doesn’t translate to more hiring.

This is a bottomline increase mission aka cost reduction. A fraction of the cost saved is the affordability for Voice AI for the enterprise.

Voice AI’s task here - Move away from Earning a fraction of the savings (difficult for India wages) to a Digitisation Budget.

2) Revenue Augmentation or Infinite Supply missions - These are use-cases where low CPA leads can be acquired via marketing and any number of conversions would be better for the topline.

An obvious example is a real estate company that wants to call dormant leads that once expressed interest or a 2 sided marketplace that wants to increase fleet operations or last mile delivery partners in a certain market.

The general principle here is to own the outcome as best possible and negotiate on ROI versus headline price.


Evolution of different flavours of Voice AI Businesses

The four Voice AI businesses mapped across 4 quadrants

These are not merely four pricing models. They are four different businesses. A business might choose to play dominantly in one quadrant and across even two or three quadrants, but it is a strategic choice.

The strategic choice to focus largely on one quadrant would serve the firm well and give it replicability of product and GTM.

Quadrant 1: Own the Outcome Model - AI enabled Services with a Human in the Loop.

In a revenue-growth mission, one way to escape minute-based pricing is to take responsibility for the result. The company may combine AI outreach, workflow automation and human intervention to deliver a qualified lead, collection, hire or completed transaction.

The customer is not purchasing an AI caller. It is purchasing a managed outcome.

Outcome ownership is not a pricing slogan. It is an operating model. Multiple VCs have underwritten the AI Services thesis.

Good news - India pays for outcomes. Actually, they only pay for outcomes.

Quadrant 2: High Volume. Price the minute—but at scale

Consider lead qualification where AI calls thousands of prospects and passes interested people to the customer’s sales team.

The Agent creates reach, but the customer still owns conversion. Humans are a bottleneck when it comes to scale, but they’re excellent at conversion once they get talking.

In this case, per-minute pricing may be entirely appropriate.

The catch is volume.

At thin margins, this model makes sense only if deployments are extremely large, implementations are standardised and account-management effort does not expand proportionately with revenue.

Capital is the moat and distribution is the proof of success. Expect the large ones to survive here.

Quadrant 3: Build the no-frills product - PLG / Narrow Product.

Standalone cost-reduction use cases—such as narrow CX workflows—can also sustain usage pricing.

But the product must behave like a product.

Customers should be able to configure, launch and monitor it with minimal assistance. The sales and implementation motion needs to be product-led, and the use case must be narrow enough to standardise.

Here, SMBs can be a credible route to profitability because the cost of serving the next customer is kept low.

Quadrant 4: Own the workflow

The other durable route is verticalization.

Instead of selling a general voice agent, the company can build a workflow product for a high-value sector or function and embed voice within it.

The product understands workflow state, conducts conversations, handles retries, coordinates human intervention and writes the output back to the system of record.

At Hunar.ai, for example, the opportunity is not merely to automate calls. Voice AI sits inside workflows covering sourcing, screening, onboarding, documentation, training, engagement and retention, while the product connects with existing HRMS / ATS / CRMs and any other communication systems like cloud telephony (human agents handoff).

The pricing can therefore move from minutes towards employees engaged, candidates processed, workflows completed or a platform subscription—often with usage limits underneath.

Verticalization alone does not guarantee good economics. The workflow must be substantially repeatable across customers. Otherwise, the company has replaced an infrastructure business with a custom-services business.

Nor does every workflow product immediately become a system of record. It normally starts as a system of action that coordinates work around an existing CRM, ATS or loan-management system. It becomes a system of record only when customers begin to rely on its persistent workflow state and history as authoritative data.

The durability is the workflow - but the true moat is ultimately becoming a new age system of record.


The durable ways out

The Voice AI market therefore has two particularly interesting routes away from commodity economics:

  1. Own the outcome / HITL Solutions : Operate the mission and accept responsibility for a measurable result. It would make sense to verticalise it to a high SDP sector or a group of adjacent sectors for GTM speed and specifc data which becomes the moat.
  2. Own the workflow: Build a vertical product in which voice is a important component.This product should ideally start becoming the system of record in a lot of the and this becomes the moat. This SOR layer becomes the moat.

The third way (riskier one imho) is to build a horizontal use-case agnostic infrastructure with moats being distribution and capital. There will be a dominant 2 with a follower sort of market structure which emerges after a wave of consolidation potentially.

The dangerous middle is a generic agent built on rented infrastructure, sold cheaply, with extensive custom solutioning bundled in for free.


The real pricing question

₹1.50 per minute may be a perfectly sensible price. But it is only sensible for a business designed to operate at ₹1.50 per minute.

Price the minute when you provide conversational capacity.

Price the workflow when you own the process.

Price the outcome when you own the mission.

The mistake is not charging ₹1.50.

The mistake is charging ₹1.50 for a business that requires ₹15 of human effort wrapped around it.


About Gregory Ipe

Gregory Ipe is Head of Business at Hunar.AI, where he works at the intersection of business, technology and the practical application of AI to frontline workforce challenges.

His professional background spans business leadership, B2B, e-commerce, sales and customer management, giving him a commercial perspective on how emerging technologies translate into operating models and measurable business outcomes.

At Hunar.AI, Gregory is focused on the business side of Voice AI and its application to workforce management. His perspective on pricing, deployment models and the economics of Voice AI reflects the broader question facing the industry: how conversational AI moves from an infrastructure capability into sustainable, outcome-driven businesses.


About Hunar.AI

Hunar.AI is an AI technology company focused on helping organisations hire, onboard, train, engage and retain frontline workers through AI-powered conversations.

Its platform combines Voice AI with workforce management workflows, supporting use cases across sourcing, screening, assessments, interviews, onboarding, training, engagement and retention. Hunar.AI's technology is designed for real-world conversations involving multiple languages, dialects, pauses, interruptions, noise and other complexities of frontline communication.

The company's broader platform connects AI-driven workforce workflows with existing HR, CRM, communication and data systems, positioning Voice AI as part of a larger operational workflow rather than simply a calling layer.

Hunar.AI says its mission is to help organisations hire, onboard and manage frontline workforces more effectively, with a long-term focus on building AI-native tools for a workforce that is often underserved by traditional workplace technology.


Source & Credits

This article was written by Gregory Ipe, Head of Business at Hunar.AI, and is published on KARV Tech Insider with author permission.