Comulytic Note Pro Transcription Accuracy Test and Evaluation Guide Methods

Affiliate disclosure: Purchases made through links on this page may earn this site a commission at no extra cost to you. All affiliate links carry rel="nofollow sponsored".
Comulytic Note Pro Transcription Accuracy Test Method

Direct answer: Comulytic advertises transcription accuracy of up to 98%, but that percentage should not be treated as a universal result. Accuracy changes with language, accent, room noise, microphone distance, overlapping speech, names and specialist vocabulary. The most reliable evaluation uses a human-verified reference transcript, calculates word error rate, and separately scores speakers, numbers, negations, action items and summary faithfulness.

No independent hands-on score is invented in this guide. Instead, it provides a repeatable protocol that an editor, buyer or reviewer can run with the actual device. The final result should report the test files, conditions, firmware, app version and scoring rules so another person can reproduce it.

Vendor claim: The official product page says Comulytic reaches “up to 98%” accuracy, supports 113 languages and uses specialist knowledge for fields such as finance, insurance and real estate. These are manufacturer statements, not results produced by this publication.

Check Comulytic Note Pro Availability

What does 98% accuracy mean?

An accuracy percentage is incomplete without a denominator and test method. A 98% word-level result would imply about two word errors per 100 reference words. A ten-minute business discussion can contain 1,300 to 1,700 words, so even 98% could produce 26 to 34 errors. Some errors are harmless punctuation differences. Others can reverse meaning.

Consider these examples:

Reference Incorrect transcript Risk
“Do not renew the contract” “Do renew the contract” Negation reversed
“Revenue was $15.4 million” “Revenue was $50.4 million” Material number error
“Give the account to Priya” “Give the account to Bria” Wrong person
“The dose is 0.5 milligrams” “The dose is 5 milligrams” Severe safety risk
“Ship on 14 August” “Ship on 40 August” Invalid date

Accuracy should therefore be split into ordinary words and high-consequence facts. A smooth summary can still be unsafe if it changes a number, owner or deadline.

Comu recorder in use

Use word error rate

NIST evaluates automatic speech recognition with word error rate, commonly written as WER:

WER = (Substitutions + Deletions + Insertions) ÷ Reference words

If a 1,000-word reference has 30 substitutions, 10 deletions and 10 insertions:

WER = (30 + 10 + 10) ÷ 1,000 = 5%

An approximate word accuracy can be reported as:

Word accuracy = 100% − WER

That example would produce about 95% word accuracy. WER can exceed 100% when a system inserts many extra words, which is another reason not to treat the converted percentage as a perfect measure.

NIST's OpenASR documentation confirms that deletions, insertions and substitutions form the standard calculation. A reviewer can use sclite, JiWER or another documented tool, but should keep normalisation rules consistent.

Build a representative test set

One quiet dictation is not enough. Comulytic is sold for meetings, calls, interviews and specialist work, so the test set should cover those conditions.

Test file Length Speakers Environment Main question
Clean dictation 5 minutes 1 Quiet room Baseline word recognition
Two-person meeting 15 minutes 2 Office Speaker labels and dialogue
Group meeting 20 minutes 6 Conference room Distance and attribution
Phone call 10 minutes 2 VPU mode Both call sides
Noisy café 10 minutes 2 Controlled 65dB background Noise resilience
Accented English 10 minutes 2 Quiet room Accent robustness
Specialist vocabulary 8 minutes 2 Quiet room Names, jargon and acronyms
Overlapping speech 5 minutes 3 Scripted interruptions Crosstalk handling
Numbers and dates 5 minutes 1 Quiet room Numeric fidelity
Non-English session 10 minutes 2 Quiet room Language-specific performance

This creates 93 minutes of audio. A publication can extend it, but the same files and placement should be reused when comparing Comulytic with Plaud, Genspark, TicNote or BOYA.

Control the recording setup

Document the following before every test:

  • Comulytic firmware version;
  • mobile app version;
  • phone model and operating system;
  • selected recording mode;
  • device position and orientation;
  • distance from every speaker;
  • measured background sound level;
  • language and accent;
  • internet connection during processing;
  • summary template;
  • custom vocabulary settings;
  • time from stop to completed transcript.

For comparison tests, place devices as close together as possible without covering microphones. Rotate positions between sessions to reduce seat advantage. Use the same speakers and script.

The device lists a five-metre pickup range. That is a maximum capture claim, not proof of equally accurate transcription at every point. Test one, three and five metres with the same sentence list.

Comu recorder in use

Create a human reference transcript

The reference must be more accurate than the system being evaluated. A practical process uses:

  1. a trained human transcriber;
  2. a second reviewer listening to the original audio;
  3. a final adjudication of disagreements;
  4. fixed spelling for names and brands;
  5. consistent treatment of fillers and partial words.

The rules should state if contractions are expanded, numbers are written as words, fillers are counted and punctuation is ignored. Because ordinary WER usually ignores punctuation and letter case, punctuation quality should receive a separate score.

Do not silently edit the Comulytic output before scoring. Preserve the raw export, then create a corrected copy for editorial analysis.

Score more than words

1. Word error rate

Calculate total WER and WER per test condition. An overall average can hide serious weakness in noise or group meetings.

2. Speaker-attribution error

Count utterances assigned to the wrong speaker:

Speaker error rate = Misattributed utterances ÷ Total attributed utterances

Also report missed speaker changes and unnecessary speaker splits.

3. Entity accuracy

Create a list of names, companies, products, places and acronyms in the reference. Score exact matches:

Entity recall = Correct entities ÷ Reference entities

4. Number accuracy

Score currencies, percentages, dates, times, measurements and phone numbers. A number is correct only if its value and unit match.

5. Negation accuracy

Check every “not,” “never,” “cannot,” “without” and equivalent phrase. Negation errors deserve separate disclosure even when total WER is low.

6. Punctuation and readability

Use a human rubric from one to five for sentence boundaries, paragraphs and punctuation. Do not mix this subjective score into WER.

7. Processing time

Measure from stopping the recording to receiving the complete transcript and summary. Report transfer and processing separately where possible.

Test the summary, not only the transcript

An AI summary can correct minor transcript noise through context, but it can also invent facts or omit important qualifiers.

Build a fact list from each reference session:

  • decisions;
  • owners;
  • deadlines;
  • amounts;
  • objections;
  • risks;
  • unresolved questions;
  • next actions.

Then score:

Metric Calculation
Fact recall Correct reference facts included ÷ total reference facts
Fact precision Correct facts included ÷ all summary claims
Action-item recall Correct actions included ÷ reference actions
Owner accuracy Actions assigned to correct owner ÷ actions with owners
Deadline accuracy Correct deadlines ÷ reference deadlines
Hallucination rate Unsupported claims ÷ all summary claims

A useful summary should preserve meaning and make follow-up easier. Length and confident language do not prove quality.

Test specialist vocabulary

Comulytic promotes a vertical knowledge base for finance, insurance and real estate. That claim deserves a dedicated challenge set.

Finance list

  • annualised yield;
  • EBITDA;
  • basis points;
  • loan-to-value ratio;
  • amortisation;
  • beneficiary;
  • ISIN;
  • counterparty.

Insurance list

  • comprehensive cover;
  • deductible;
  • underwriting;
  • indemnity;
  • subrogation;
  • policyholder;
  • loss adjuster;
  • riders.

Real-estate list

  • cap rate;
  • escrow;
  • encumbrance;
  • conveyancing;
  • comparable sale;
  • earnest money;
  • leasehold;
  • zoning variance.

Add local client names, product names and internal acronyms through any supported vocabulary feature. Run the test once before and once after customisation. Report the improvement rather than assuming the feature works.

Language and accent testing

The product lists 113 languages. That does not mean equal performance across all 113. Speech-recognition systems usually have different training data, dialect coverage and punctuation quality by language.

For each important language:

  1. use native speakers;
  2. create a human reference in the same language;
  3. include formal and conversational speech;
  4. include regional names;
  5. test code-switching;
  6. score the original-language transcript;
  7. score translation separately if used.

Translation quality must not be confused with transcription quality. A correct transcript can be translated badly, and an incorrect transcript can produce a fluent but wrong translation.

Noise and distance testing

Use controlled, repeatable background audio through a calibrated speaker. Suggested levels are:

Environment Approximate controlled level Purpose
Quiet office 35 to 40dB Baseline
Open office 50 to 55dB Everyday work
Busy café 60 to 70dB Difficult conversation
Vehicle cabin Varies Road and engine noise

Sound-level readings depend on equipment and position, so report the measuring device. Never compare a quiet-room Comulytic transcript with a café transcript from a competitor.

At each noise level, test speakers at one, three and five metres. The result should show a matrix, not one blended percentage.

Phone-call testing

The VPU records vibrations transmitted through a phone. Performance can change with:

  • phone model;
  • case material and thickness;
  • recorder placement;
  • speaker volume;
  • VoIP versus cellular call;
  • handset versus speakerphone;
  • network compression.

Record both sides with informed consent. Confirm that exported audio contains both speakers and that the transcript labels them consistently. The call test should use the same phone and case when comparing devices.

Compare free and Premium output

Comulytic says core transcription and basic summaries are free, while Premium adds richer templates and analysis. The raw transcript should be tested on the free entitlement first. Then compare:

  • basic summary;
  • Premium summary;
  • action-item extraction;
  • AI advice;
  • client-profile updates;
  • question answering.

This reveals if Premium improves factual output or mainly changes presentation. A paid feature is valuable only if it saves correction time or creates a useful workflow result.

Accuracy scorecard template

Test WER Speaker error Entity recall Number accuracy Summary fact recall Hallucinations
Clean dictation TBD N/A TBD TBD TBD TBD
Two-person meeting TBD TBD TBD TBD TBD TBD
Group meeting TBD TBD TBD TBD TBD TBD
Phone call TBD TBD TBD TBD TBD TBD
Noisy café TBD TBD TBD TBD TBD TBD
Specialist vocabulary TBD TBD TBD TBD TBD TBD

TBD is intentional. It prevents a research framework from being misrepresented as completed product testing. Editors should replace every cell only after preserving the source files and calculation output.

How to interpret results

The best score depends on use:

  • A journalist may prioritise verbatim words, names and searchable time codes.
  • A salesperson may prioritise objections, commitments and follow-up actions.
  • A student may prioritise coverage, headings and definitions.
  • A lawyer may require exact wording and manual verification.
  • A clinician must follow organisational policy and verify medical facts.

No AI transcript should become an unquestioned source of truth in a high-consequence workflow. The original audio should remain available, and a human should verify critical passages.

Market relevance

The speech and voice-recognition market is expanding, but published estimates define it differently. MarketsandMarkets valued its category at $9.66 billion in 2025 and projected $23.11 billion by 2030. Growth creates more products and faster feature releases, but also more unqualified accuracy marketing.

An evidence-based review can outperform a generic listicle by publishing:

  • downloadable test scripts;
  • recording conditions;
  • raw and reference transcripts where consent allows;
  • scoring code;
  • per-condition results;
  • firmware and app versions;
  • correction-time measurements.

That evidence is also useful for Google AI Overviews because it provides clear definitions, methods and attributable conclusions.

Final assessment

Comulytic's 98% claim is plausible only as an “up to” figure under unspecified favourable conditions. It is not enough to predict performance in a buyer's meeting room.

The product should be judged through WER, speaker accuracy, entity accuracy, number fidelity, negations, summary facts and correction time. A strong result in clean English should not conceal weakness in noisy, multilingual or overlapping speech.

Comulytic has a meaningful advantage for testing because unlimited core transcription allows repeated trials without consuming a small monthly allowance. That commercial advantage is real even before an accuracy winner is known.

View Comulytic Note Pro Details
Quick answers

Frequently asked questions

Clear answers to common questions readers check before choosing a Comu recorder.

Is Comulytic Note Pro really 98% accurate?

Comulytic says “up to 98%.” The result will vary by language, noise, distance, speaker overlap and vocabulary. Independent conditions are not published with the claim.

What is word error rate?

WER divides substitutions, deletions and insertions by the number of words in a human reference transcript. Lower is better.

Is 98% accuracy the same as 2% WER?

It can be presented that way in a simplified conversion, but WER has complications and can exceed 100%. The scoring rules must be disclosed.

How should speaker identification be tested?

Count utterances assigned to the wrong person, missed speaker changes and unnecessary speaker splits in a multi-person reference session.

Does Comulytic support specialist terms?

The company promotes vertical vocabulary for finance, insurance and real estate. A reviewer should test a fixed term list before and after custom-vocabulary setup.

Can one quiet recording prove accuracy?

No. A representative test needs quiet, noisy, distant, phone, group, accented, multilingual and overlapping-speech conditions.

Should punctuation count in WER?

Conventional WER often normalises punctuation and case. Punctuation should receive a separate readability score.

How should AI summaries be scored?

Score fact recall, precision, action items, owners, deadlines and unsupported claims against a human-created fact list.

Does Premium improve transcription accuracy?

That should be tested. Premium may mainly affect templates and analysis, so free and paid outputs should be compared using identical audio.

Can Comulytic be trusted for legal or medical records?

Critical content requires human verification, organisational approval and access to original audio. No consumer accuracy claim removes that responsibility.

What is the most dangerous transcript error?

Negations, amounts, dates, doses, names and assigned owners can change decisions. They deserve separate scoring and manual review.

How can a publisher make the test credible?

Publish the script, conditions, versions, raw output, reference transcript, scoring rules and calculation results, subject to consent and privacy limits.