Invoca formats spoken numbers, currency, and URLs the way they’d appear in print, and phrase spotting matches flexibly against that formatted text.
Transcription
Transcripts format spoken phrases — numbers, currency, time, percentages, URLs, hyphenated words, brands, and slashes — the way they’d appear in print, rather than word for word.
Transcription details: pauses in read-outs of numbers can split them in the transcript. Number pronunciation variance — for example, “4050” read as “Forty Fifty” or “Four Oh Five Oh” — can also cause inaccuracies.
Phrase spotting
Phrase spotting formats text as it would appear in print — for example, matching against “$100” rather than “one hundred dollars.” Invoca first matches words as transcribed, then reverses the formatting and retries, for more accurate matching. The UI accommodates special characters:Flexible matching
- Word stemming: “walk” matches “walked” or “walking,” and vice versa.
- Number partial matches: a phrase that’s just a number, such as “100,” matches “$100” or “100%” in the transcript — but not the reverse. A phrase with a number plus a symbol only matches the exact same number and symbol.
- Slashes and dashes: treated as spaces when comparing. A phrase like “this-that” matches “this-that,” “this that,” or “this/that” in the transcript.