Bengali Transcription Expert
$10-20
/hr
Paid work turning recorded speech into text a speech model can learn from: verbatim transcription with timestamps, speaker and noise tags, correcting machine output, and grading what a speech AI produced. Almost every listing names a language and often a region. Voice recording, where you are the speaker, is listed separately on voice and audio. It is work-from-home work by nature: headphones, a quiet room and the platform's editor. Pay on this page runs $11 to $23/hr for the middle half of roles.
Part of: Creative & Design
A speech model learns from pairs of audio and text, and the text has to match the audio exactly, not what the speaker meant to say. So the work on this page is verbatim: every false start, repeated word, "um", half-finished sentence and overlapping speaker, marked the way the project's guideline says. Clean transcription, the kind a journalist wants, is a different product and usually a different listing.
Four task types cover most listings. Verbatim transcription from scratch, with timestamps. Speaker diarization, labeling who is talking and where turns start and end. Correction of machine output, where an automatic speech recognition (ASR) system has produced a draft and you fix it, which is faster per minute and is becoming the common form. And speech AI evaluation, where you listen to what a speech model produced or transcribed and grade it. Listings also cover document transcription (reading handwritten or PDF text into a structured format), which uses the same conventions without the audio.
The language in the title decides whether you are eligible. Almost every listing names a language and often a region (Vietnamese as spoken in Taiwan, Portuguese as spoken in Brazil), because dialect and accent are what the model is being trained to handle. Voice recording, where you are the speaker rather than the transcriber, is listed separately on voice and audio.
An original example, written for this page. The clip is ninety seconds of two people arranging a delivery over the phone. There is a dog barking in the background, the speakers talk over each other twice, one of them reads a reference number as "double oh seven", and the call drops for three seconds.
Your guideline says: verbatim, tag non-speech events in square brackets, mark overlap with a defined symbol, write numbers as spoken, and never guess an inaudible word. The output is text with speaker labels and timestamps.
| What you hear | What you write |
|---|---|
| Speaker 1 starts the address, Speaker 2 talks over the street name | Both turns, each tagged with the overlap marker at the words that collide |
| "double oh seven" | double oh seven, as spoken, with a number tag if the guideline uses one |
| Dog barks over "Tuesday" | [dog barking] and the word, or [inaudible] if the word cannot be recovered |
| Three seconds of silence after the drop | A silence tag with the duration, not a guess at the missing words |
The thing that fails people is guessing. A plausible word written over an inaudible one looks fine to a human reader and is poison to a model, and quality review compares your text against the audio word by word. The second thing is inconsistency: the same event tagged two ways in one file.
You need native or near-native command of the listed language and its regional variety, typing speed that makes verbatim work pay, and the patience to apply a long guideline the same way for hours. Transcriptionists, subtitlers, linguists and language teachers do well, as do bilingual people who have done any kind of annotation work. Clinical and legal transcription experience helps with the specialist listings.
Speech AI evaluation and the engineering-flavored listings (ASR, speech data linguistics) want more: a linguistics or speech technology background, sometimes a degree, and comfort with the tooling. Those are a minority of the roles here and pay at the top of the page's range.
It does not suit people who want the highest hourly rate on the board. This page's pay block is usually the lowest of the four kind-of-work hubs, because verbatim transcription is priced near other language work rather than near expert evaluation. Where it pays well is per-minute rates in a less common language, where a fast, accurate transcriptionist with the right dialect has little competition.
Expect a short audio sample to transcribe under the project's rules, scored on word accuracy and on whether you used the tags correctly, and a language or dialect check that may be a conversation. Listings on this page rarely describe an interview; where they do, it is a review of your delivered sample. Equipment questions are simple: headphones, a quiet room and a reliable connection are the usual requirements, and foot pedals are not expected.
Before the test, read the guideline twice and transcribe a minute of any recording the strict way, tagging everything. The Academy module Your First Annotation Task is about text rating rather than audio, but its lesson on applying a guideline consistently is the one transcription tests measure, and How AI Training Jobs Work explains how quality review and queues work on these platforms.
The pay block at the top of this page is computed from the open listings below it: typical (median) hourly rate, the range most roles pay, and how many posted rates it is based on, updated on the date shown. The method is on the pay methodology page.
Listed rates on this page currently run $11 to $23/hr for the middle half of roles.
Fewer listings here state a rate than on the other hubs, and some pay per audio minute or per task rather than per hour; those are excluded from the hourly figure, so read the listing for its own terms. The notes beside this section say when our own estimates or one platform dominate the numbers. Speech AI evaluation and linguistics roles sit at the top of the range; verbatim transcription in a widely spoken language at the bottom.
Related reading: voice and audio AI training jobs covers the recording side, and one worker's voice recording experience shows what a speech data project is like from the inside. Adjacent kinds of work: RLHF evaluation grades text, red teaming tests refusals, and agentic tasks cover agents that act.
76 Jobs
Page 1 of 3

Trending jobs + new guides + real pay data. Every Tuesday.
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
$10-35
/hr
One browser alert a day when new roles land, Transcription & Speech Data first. No email needed.
The listings on this page currently state $11 to $23/hr for the middle half of roles, with the median and the number of rates behind it in the pay block at the top. Some roles pay per audio minute or per task instead of per hour; those are excluded from the hourly figure, so check each listing's own terms.
Usually no. Platforms test you on a short audio sample under their guideline and on your command of the language and dialect. Experience with verbatim conventions, subtitling or any annotation work helps you pass faster; a typing speed that makes verbatim work pay is the practical requirement.
In transcription you listen and write; in voice recording you are the speaker and the recording is the product. They train different parts of a speech system and are listed separately here. Voice recording and voice acting roles are on the voice and audio page.
Diarization is labeling who is speaking and when: marking each turn with a speaker label and the time it starts and ends, including overlaps. It is a task on its own in some listings and part of verbatim transcription in others.
The search and platform filters above the list show today's languages. Demand follows which languages a lab is shipping speech features in; on October 1, 2026 the open roles were mostly Nordic languages, Indian and Southeast Asian languages, and regional variants of Spanish, Portuguese and Chinese. Less common languages pay more per minute because fewer transcriptionists qualify.
Both, and it varies by listing. Per-hour contracts are the majority on this page; per-audio-minute and per-task rates appear on larger batch projects. Per-minute pay rewards speed, so a rate that looks low per minute can beat an hourly rate for a fast transcriptionist, and the reverse.
Closed headphones, a quiet room, a reliable internet connection and a computer that runs the platform's browser tool. Foot pedals and professional transcription software are not required; projects supply the tool and the guideline.
This page is showing 76 listings, sourced from 6 platforms: Alignerr (18), rws (17), Micro1 (15), Turing (11) and Mercor (10). The most recent one was picked up 2 days ago. The count moves as platforms post and close roles, so treat it as a snapshot of today rather than a fixed figure.
Browse all job categories and find your next AI training opportunity.
View All Categories