
Direct answer: Comulytic advertises transcription accuracy of up to 98%, but that percentage should not be treated as a universal result. Accuracy changes with language, accent, room noise, microphone distance, overlapping speech, names and specialist vocabulary. The most reliable evaluation uses a human-verified reference transcript, calculates word error rate, and separately scores speakers, numbers, negations, action items and summary faithfulness.
No independent hands-on score is invented in this guide. Instead, it provides a repeatable protocol that an editor, buyer or reviewer can run with the actual device. The final result should report the test files, conditions, firmware, app version and scoring rules so another person can reproduce it.
Vendor claim: The official product page says Comulytic reaches “up to 98%” accuracy, supports 113 languages and uses specialist knowledge for fields such as finance, insurance and real estate. These are manufacturer statements, not results produced by this publication.
What does 98% accuracy mean?
An accuracy percentage is incomplete without a denominator and test method. A 98% word-level result would imply about two word errors per 100 reference words. A ten-minute business discussion can contain 1,300 to 1,700 words, so even 98% could produce 26 to 34 errors. Some errors are harmless punctuation differences. Others can reverse meaning.
Consider these examples:
| Reference | Incorrect transcript | Risk |
|---|---|---|
| “Do not renew the contract” | “Do renew the contract” | Negation reversed |
| “Revenue was $15.4 million” | “Revenue was $50.4 million” | Material number error |
| “Give the account to Priya” | “Give the account to Bria” | Wrong person |
| “The dose is 0.5 milligrams” | “The dose is 5 milligrams” | Severe safety risk |
| “Ship on 14 August” | “Ship on 40 August” | Invalid date |
Accuracy should therefore be split into ordinary words and high-consequence facts. A smooth summary can still be unsafe if it changes a number, owner or deadline.

Use word error rate
NIST evaluates automatic speech recognition with word error rate, commonly written as WER:
WER = (Substitutions + Deletions + Insertions) ÷ Reference words
If a 1,000-word reference has 30 substitutions, 10 deletions and 10 insertions:
WER = (30 + 10 + 10) ÷ 1,000 = 5%
An approximate word accuracy can be reported as:
Word accuracy = 100% − WER
That example would produce about 95% word accuracy. WER can exceed 100% when a system inserts many extra words, which is another reason not to treat the converted percentage as a perfect measure.
NIST's OpenASR documentation confirms that deletions, insertions and substitutions form the standard calculation. A reviewer can use sclite, JiWER or another documented tool, but should keep normalisation rules consistent.
Build a representative test set
One quiet dictation is not enough. Comulytic is sold for meetings, calls, interviews and specialist work, so the test set should cover those conditions.
| Test file | Length | Speakers | Environment | Main question |
|---|---|---|---|---|
| Clean dictation | 5 minutes | 1 | Quiet room | Baseline word recognition |
| Two-person meeting | 15 minutes | 2 | Office | Speaker labels and dialogue |
| Group meeting | 20 minutes | 6 | Conference room | Distance and attribution |
| Phone call | 10 minutes | 2 | VPU mode | Both call sides |
| Noisy café | 10 minutes | 2 | Controlled 65dB background | Noise resilience |
| Accented English | 10 minutes | 2 | Quiet room | Accent robustness |
| Specialist vocabulary | 8 minutes | 2 | Quiet room | Names, jargon and acronyms |
| Overlapping speech | 5 minutes | 3 | Scripted interruptions | Crosstalk handling |
| Numbers and dates | 5 minutes | 1 | Quiet room | Numeric fidelity |
| Non-English session | 10 minutes | 2 | Quiet room | Language-specific performance |
This creates 93 minutes of audio. A publication can extend it, but the same files and placement should be reused when comparing Comulytic with Plaud, Genspark, TicNote or BOYA.
Control the recording setup
Document the following before every test:
- Comulytic firmware version;
- mobile app version;
- phone model and operating system;
- selected recording mode;
- device position and orientation;
- distance from every speaker;
- measured background sound level;
- language and accent;
- internet connection during processing;
- summary template;
- custom vocabulary settings;
- time from stop to completed transcript.
For comparison tests, place devices as close together as possible without covering microphones. Rotate positions between sessions to reduce seat advantage. Use the same speakers and script.
The device lists a five-metre pickup range. That is a maximum capture claim, not proof of equally accurate transcription at every point. Test one, three and five metres with the same sentence list.

Create a human reference transcript
The reference must be more accurate than the system being evaluated. A practical process uses:
- a trained human transcriber;
- a second reviewer listening to the original audio;
- a final adjudication of disagreements;
- fixed spelling for names and brands;
- consistent treatment of fillers and partial words.
The rules should state if contractions are expanded, numbers are written as words, fillers are counted and punctuation is ignored. Because ordinary WER usually ignores punctuation and letter case, punctuation quality should receive a separate score.
Do not silently edit the Comulytic output before scoring. Preserve the raw export, then create a corrected copy for editorial analysis.
Score more than words
1. Word error rate
Calculate total WER and WER per test condition. An overall average can hide serious weakness in noise or group meetings.
2. Speaker-attribution error
Count utterances assigned to the wrong speaker:
Speaker error rate = Misattributed utterances ÷ Total attributed utterances
Also report missed speaker changes and unnecessary speaker splits.
3. Entity accuracy
Create a list of names, companies, products, places and acronyms in the reference. Score exact matches:
Entity recall = Correct entities ÷ Reference entities
4. Number accuracy
Score currencies, percentages, dates, times, measurements and phone numbers. A number is correct only if its value and unit match.
5. Negation accuracy
Check every “not,” “never,” “cannot,” “without” and equivalent phrase. Negation errors deserve separate disclosure even when total WER is low.
6. Punctuation and readability
Use a human rubric from one to five for sentence boundaries, paragraphs and punctuation. Do not mix this subjective score into WER.
7. Processing time
Measure from stopping the recording to receiving the complete transcript and summary. Report transfer and processing separately where possible.
Test the summary, not only the transcript
An AI summary can correct minor transcript noise through context, but it can also invent facts or omit important qualifiers.
Build a fact list from each reference session:
- decisions;
- owners;
- deadlines;
- amounts;
- objections;
- risks;
- unresolved questions;
- next actions.
Then score:
| Metric | Calculation |
|---|---|
| Fact recall | Correct reference facts included ÷ total reference facts |
| Fact precision | Correct facts included ÷ all summary claims |
| Action-item recall | Correct actions included ÷ reference actions |
| Owner accuracy | Actions assigned to correct owner ÷ actions with owners |
| Deadline accuracy | Correct deadlines ÷ reference deadlines |
| Hallucination rate | Unsupported claims ÷ all summary claims |
A useful summary should preserve meaning and make follow-up easier. Length and confident language do not prove quality.
Test specialist vocabulary
Comulytic promotes a vertical knowledge base for finance, insurance and real estate. That claim deserves a dedicated challenge set.
Finance list
- annualised yield;
- EBITDA;
- basis points;
- loan-to-value ratio;
- amortisation;
- beneficiary;
- ISIN;
- counterparty.
Insurance list
- comprehensive cover;
- deductible;
- underwriting;
- indemnity;
- subrogation;
- policyholder;
- loss adjuster;
- riders.
Real-estate list
- cap rate;
- escrow;
- encumbrance;
- conveyancing;
- comparable sale;
- earnest money;
- leasehold;
- zoning variance.
Add local client names, product names and internal acronyms through any supported vocabulary feature. Run the test once before and once after customisation. Report the improvement rather than assuming the feature works.
Language and accent testing
The product lists 113 languages. That does not mean equal performance across all 113. Speech-recognition systems usually have different training data, dialect coverage and punctuation quality by language.
For each important language:
- use native speakers;
- create a human reference in the same language;
- include formal and conversational speech;
- include regional names;
- test code-switching;
- score the original-language transcript;
- score translation separately if used.
Translation quality must not be confused with transcription quality. A correct transcript can be translated badly, and an incorrect transcript can produce a fluent but wrong translation.
Noise and distance testing
Use controlled, repeatable background audio through a calibrated speaker. Suggested levels are:
| Environment | Approximate controlled level | Purpose |
|---|---|---|
| Quiet office | 35 to 40dB | Baseline |
| Open office | 50 to 55dB | Everyday work |
| Busy café | 60 to 70dB | Difficult conversation |
| Vehicle cabin | Varies | Road and engine noise |
Sound-level readings depend on equipment and position, so report the measuring device. Never compare a quiet-room Comulytic transcript with a café transcript from a competitor.
At each noise level, test speakers at one, three and five metres. The result should show a matrix, not one blended percentage.
Phone-call testing
The VPU records vibrations transmitted through a phone. Performance can change with:
- phone model;
- case material and thickness;
- recorder placement;
- speaker volume;
- VoIP versus cellular call;
- handset versus speakerphone;
- network compression.
Record both sides with informed consent. Confirm that exported audio contains both speakers and that the transcript labels them consistently. The call test should use the same phone and case when comparing devices.
Compare free and Premium output
Comulytic says core transcription and basic summaries are free, while Premium adds richer templates and analysis. The raw transcript should be tested on the free entitlement first. Then compare:
- basic summary;
- Premium summary;
- action-item extraction;
- AI advice;
- client-profile updates;
- question answering.
This reveals if Premium improves factual output or mainly changes presentation. A paid feature is valuable only if it saves correction time or creates a useful workflow result.
Accuracy scorecard template
| Test | WER | Speaker error | Entity recall | Number accuracy | Summary fact recall | Hallucinations |
|---|---|---|---|---|---|---|
| Clean dictation | TBD | N/A | TBD | TBD | TBD | TBD |
| Two-person meeting | TBD | TBD | TBD | TBD | TBD | TBD |
| Group meeting | TBD | TBD | TBD | TBD | TBD | TBD |
| Phone call | TBD | TBD | TBD | TBD | TBD | TBD |
| Noisy café | TBD | TBD | TBD | TBD | TBD | TBD |
| Specialist vocabulary | TBD | TBD | TBD | TBD | TBD | TBD |
TBD is intentional. It prevents a research framework from being misrepresented as completed product testing. Editors should replace every cell only after preserving the source files and calculation output.
How to interpret results
The best score depends on use:
- A journalist may prioritise verbatim words, names and searchable time codes.
- A salesperson may prioritise objections, commitments and follow-up actions.
- A student may prioritise coverage, headings and definitions.
- A lawyer may require exact wording and manual verification.
- A clinician must follow organisational policy and verify medical facts.
No AI transcript should become an unquestioned source of truth in a high-consequence workflow. The original audio should remain available, and a human should verify critical passages.
Market relevance
The speech and voice-recognition market is expanding, but published estimates define it differently. MarketsandMarkets valued its category at $9.66 billion in 2025 and projected $23.11 billion by 2030. Growth creates more products and faster feature releases, but also more unqualified accuracy marketing.
An evidence-based review can outperform a generic listicle by publishing:
- downloadable test scripts;
- recording conditions;
- raw and reference transcripts where consent allows;
- scoring code;
- per-condition results;
- firmware and app versions;
- correction-time measurements.
That evidence is also useful for Google AI Overviews because it provides clear definitions, methods and attributable conclusions.
Final assessment
Comulytic's 98% claim is plausible only as an “up to” figure under unspecified favourable conditions. It is not enough to predict performance in a buyer's meeting room.
The product should be judged through WER, speaker accuracy, entity accuracy, number fidelity, negations, summary facts and correction time. A strong result in clean English should not conceal weakness in noisy, multilingual or overlapping speech.
Comulytic has a meaningful advantage for testing because unlimited core transcription allows repeated trials without consuming a small monthly allowance. That commercial advantage is real even before an accuracy winner is known.
Frequently asked questions
Clear answers to common questions readers check before choosing a Comu recorder.
Is Comulytic Note Pro really 98% accurate?
Comulytic says “up to 98%.” The result will vary by language, noise, distance, speaker overlap and vocabulary. Independent conditions are not published with the claim.
What is word error rate?
WER divides substitutions, deletions and insertions by the number of words in a human reference transcript. Lower is better.
Is 98% accuracy the same as 2% WER?
It can be presented that way in a simplified conversion, but WER has complications and can exceed 100%. The scoring rules must be disclosed.
How should speaker identification be tested?
Count utterances assigned to the wrong person, missed speaker changes and unnecessary speaker splits in a multi-person reference session.
Does Comulytic support specialist terms?
The company promotes vertical vocabulary for finance, insurance and real estate. A reviewer should test a fixed term list before and after custom-vocabulary setup.
Can one quiet recording prove accuracy?
No. A representative test needs quiet, noisy, distant, phone, group, accented, multilingual and overlapping-speech conditions.
Should punctuation count in WER?
Conventional WER often normalises punctuation and case. Punctuation should receive a separate readability score.
How should AI summaries be scored?
Score fact recall, precision, action items, owners, deadlines and unsupported claims against a human-created fact list.
Does Premium improve transcription accuracy?
That should be tested. Premium may mainly affect templates and analysis, so free and paid outputs should be compared using identical audio.
Can Comulytic be trusted for legal or medical records?
Critical content requires human verification, organisational approval and access to original audio. No consumer accuracy claim removes that responsibility.
What is the most dangerous transcript error?
Negations, amounts, dates, doses, names and assigned owners can change decisions. They deserve separate scoring and manual review.
How can a publisher make the test credible?
Publish the script, conditions, versions, raw output, reference transcript, scoring rules and calculation results, subject to consent and privacy limits.