Audio features

Hear it. Catch it. Say it back.

Poly turns real audio into connected listening, dictation, pronunciation, and speaking practice—without separating them from review.

01

Audio lessons

Import MP3, M4A, WAV, and supported audio. Continue listening with lock-screen controls and background playback.

02

Synced transcripts

Generate a transcript or import LRC, SRT, VTT, or text. Tap a line to seek; active words follow playback.

03

Sentence mining

Save useful transcript lines as normal review cards with paired meaning and source context. Audio quiz modes stay optional.

04

Listen mode

Hear target-language audio before seeing text. Source clip plays when attached; pronunciation voice supplies fallback.

05

Dictation

Replay target audio, type exactly what you hear, then check normalized text match without punctuation or accent penalties.

06

Speaking

Speak target answer, inspect recognized transcript, and compare intelligibility using target-language recognition locale.

07

Pronunciation

Hear revealed answers with best installed Apple voice. AI+ adds professional neural pronunciation after reveal, with device fallback.

08

Shadowing studio

Listen to one sentence, record yourself, replay both versions, and compare rhythm, stress, timing, and clarity.

One card, six complementary skills.

Understand checks recognition. Produce checks recall. Listen checks comprehension. Dictation checks exact hearing. Type checks written production. Speak checks intelligibility. Poly remembers chosen mode for each card.

Permissions and privacy

Speaking asks for Microphone and Speech Recognition access. Apple may process recognition on device or through its service, depending on language and device support. Shadow recordings stay temporary and clear when player closes.

What score means

Text match measures whether words were recognized, not accent quality. Use Shadowing to compare rhythm and timing. Poly ignores letter case, punctuation, and diacritics during text matching.