All case studies Case 09 · AI, founder
ai, founder 2026 to present

Accent Studio

A voice-data tool to help AI learn to speak African and underserved languages fluently. Founded, designed, built and shipped by me.

  • Shapes propositions and operating models
  • Commercial roadmaps
  • Leads discovery
  • AI and data
Where
Accent Studio (accentstudio.io)
When
2026 to present
My role
Founder, designer, builder
Themes
FounderAIInclusionTwo-sided serviceProposition designLive product

12 locales

on day 1, 4 pilot cities, 1 founder

  1. 01Situation

    More than 80% of open voice-AI training hours are in English or Mandarin. Models fumble the moment they meet a Lagos customer or a Nairobi merchant, and the data to fix that barely exists.

  2. 02The move

    I designed a 2-sided service: a voice improv game that pays native speakers properly, with consent, PII redaction and linguist verification built in, plus 4 commercial tiers so labs can sample before they commit and a "Hear the gap" demo so their own ears make the case.

  3. 03Result

    It is live at accentstudio.io, launching first in Lagos, Nairobi, Accra and Johannesburg across 12 locales. The proposition, the operating model and the product were all shaped and shipped by 1 person.

Exhibit 09

The voice-data flow, "Hear the gap", and the live product

A scene from the game, the pipeline behind it, the gap your ears can hear, and the row a lab pays for.

Part A

The consumer side: a Live Arena scene.

LIVE · ARENA

You are dealt a persona, in your language. A twist drops mid-scene. The chemistry meter fills as the scene flows.

CHEMISTRY0%

Part B

The pipeline behind every scene.

Raw recording. Dual-channel audio, 24 kHz, 1 voice per channel. Each player on their own phone, 1 take.

Part C

Hear the gap.

Same line, 2 takes. First, a state-of-the-art voice model attempting it. Then a native speaker from the studio. The words light up as they are spoken.

Nigerian Pidgin

English sourceGood morning. I woke up this morning and saw that ₦50,000 left my account through a POS I never used.

Morning o. Naso I wake this morning see say POS wey i no use commot 50k for my account!

Yoruba

English sourceGood morning. I bought this phone from your shop last week, but it stopped working.

Ẹ káàrọ̀ o. Mo ra fóònù yìí ní ilé ìtajà yin ní ọ̀sẹ̀ tó kọjá, ṣùgbọ́n kò ṣiṣẹ́ mọ́.

Part D

The live site.

Shaped, built and shipped by me.Founder, designer, builder

Part E

The row a lab pays for.

Every scene becomes turns like this one. Dual-channel WAV beside it, Parquet index over it, consent hash inside it. Each field is there because a buyer asked what they would receive.

// manifest.jsonl · 1 turn
{
}

Why this field exists

verified_text

What was actually said, verified by a certified linguist. The ground truth.

  • Dual-channel WAV · 24 kHz
  • JSONL, 1 row per session
  • Parquet turn-level index
  • Word-level alignment, optional
  • Consent log, SHA-256 anchored
  • PGP-signed bundle
Jargon, explainedTTSCorpusWhy these languages

The situation

Accent Studio is the product I am most excited about having built. It is a voice-data tool whose purpose is to help train large language models to speak fluently in African and underserved languages, the languages the big models handle badly or not at all, because the data to teach them simply does not exist in the volumes the models were trained on.

It is a business-to-business-to-consumer model. On the consumer side, we give people a gamified tool that gets them having natural, back-and-forth conversations in their own language. We pay them for their time, properly, at rates calibrated to the hour. On the business side, we take those recordings and structure them into something AI labs can consume plug-and-play, as training data that is actually usable rather than raw and messy.

Pilot cities are Lagos, Nairobi, Accra and Johannesburg. The pricing is tiered: pilot corpus, standing subscription, custom commission, aligned corpus.

Why it matters to me

This sits at the exact intersection of the things I care about: AI, inclusion, and the languages and people that technology routinely leaves behind. The whole of my career has a thread of designing for the people the system fails first, and Accent Studio is that thread pointed directly at the future of AI. If the models only ever learn the languages that are already well represented, the exclusion that already exists just gets baked in permanently. This is my small attempt to push against that.

I designed and built it end to end on my own, from framing the problem through to a live product.

What I did

I shaped the proposition first. 3 distinct tiers to let labs sample before committing: pilot corpus for a one-off scoped dataset, standing subscription for ongoing fresh verified hours, custom commission for buyer-authored scenarios. Later I added an aligned corpus tier at a premium price, which solved a specific objection labs had around alignment to English source text.

I designed the two-sided service: consent, PII redaction, quality verification by linguists, payouts through M-Pesa, Paystack, Stripe, PayPal and USDC. The consumer tool is a gamified improv game: Ping-Pong for async play, Live Arena for real-time two-player scenes with a chemistry meter. So the data collection feels like play, not labour.

And I built the business case for the labs honestly: a side-by-side demo called “Hear the gap” comparing a state-of-the-art text-to-speech attempt against a native-speaker recording in Nigerian Pidgin and Yoruba, so a lab’s own ears make the case before any pitch deck does.

What it proves

More than the technical build, Accent Studio proves I can shape a proposition, test it with real buyers, build the operating model, and ship a live product. Those are the things organisations ask a service designer to do for the start-ups and businesses they support. Here, I am the small business.

The lesson

If AI is going to speak for everyone, it needs to hear everyone first.

And building that bridge is service design: consent, payment, verification, trust, and the gamified loop that makes contributing feel worth doing.

Research

How I found out .

Each method, and why I chose it.

  1. WhyThe objection I expected was price. The real one was alignment to English source text. That conversation created the fourth tier.

Who it was for

The people it had to work for.

Composite and anonymised. Needs and frictions kept, names removed.

Persona 1 of 4

The lab's ML lead

Training a voice-to-voice model that fumbles the moment it meets a Nigerian customer.

Needs
Clean, dual-channel, emotionally varied conversational audio with word-level alignment.
Friction
Everything available is English, scraped, or both.
What surfaced

Written on the
inside cover

  1. 80%+ of open voice-AI training hours are English or Mandarin. Nigerian Pidgin has around 400 public hours.
  2. Labs wanted to hear the gap before reading about it. The demo closes conversations a deck could not open.
  3. Contributors wanted anonymity more than money. Blind pairing became a core mechanic, not a privacy feature.
  4. Literacy must not gate participation. Every scene card is delivered as audio, with text as a fallback.
Artefacts

The things people pointed at.

Every decision in this story had something on the wall behind it.

Most of this is under NDA, so each piece is described rather than shown.

  1. 01

    Hear the gap

    TTS versus native speaker, 2 languages, same line. Part C of the exhibit, with the real audio.

  2. 02

    Turn schema

    The JSONL row a lab receives. Part E of the exhibit, every field explained.

  3. 03

    Contributor agreement v1.2

    11 pages, plain-language summary first, SHA-256 printed in full on the Labs page.

  4. 04

    Tier sheet

    Pilot corpus, standing subscription, custom commission, aligned corpus. Non-exclusive by default.

  5. 05

    Reach calculator

    Same product, same budget, with and without local-language voice data. Nigeria goes from 35M to 200M addressable.

Outcomes

What changed, and what I keep.

12 locales
live or ramping on day 1, 3 more announced
4 cities
Lagos, Nairobi, Accra, Johannesburg
4 tiers
pilot corpus, subscription, commission, aligned corpus
7 of 7
EDPB valid-consent criteria met by the contributor agreement
Impact
  1. A supply of born-digital, consented, linguist-verified conversational audio in languages where public corpora hold a few hundred hours.
  2. Contributors are paid a published rate per verified hour, on rails they already use, with their identity in a vault buyers never receive.
  3. The reach argument (Nigeria from around 35M addressable to around 200M with local-language voice) gives labs a commercial reason to care about inclusion, not just an ethical one.
Takeaways
  1. 01 If AI is going to speak for everyone, it needs to hear everyone first.
  2. 02 Consent, payment, verification and trust are the service. The game is how people are willing to give them.
  3. 03 Let buyers hear the gap before they read about it.
What I would do differently

I would have built the "Hear the gap" demo before the pricing page. Labs listened before they read, and the demo closed conversations the deck could not open.

Contact

Got a problem nobody has framed yet?

A service to fix, a programme that has stalled, or a proposition that needs shaping. I'd like to hear about it.

mubbyrobyn@gmail.com
LinkedIn
in/mubaraqrobyn ↗
CV
Mubby-Robyn-CV.pdf ↓
Based in
Reading, UK. Hybrid or remote.
Work status
Full right to work in the UK. BPSS eligible.