How we test and review

Every review, comparison, and benchmark on meetingstack follows a consistent process. This page explains what we measure, how we measure it, and where our limitations are.

Evaluation process

Transcription accuracy claims come from our own benchmark runs against ground-truth transcripts of real meeting recordings, with accents, crosstalk, background noise, and technical jargon, because that is what real meetings sound like. The rest of each review is structured desk research: vendor docs, pricing pages, changelogs, and user reports. We label which is which.

Every price comes from the vendor's public pricing page, read on the date stamped on the review. Where we have gone through a signup flow and found a different number at checkout, we report the checkout number and say so.

What we measure

Transcription accuracy

Word Error Rate (WER) measured against human-verified ground truth transcripts. We test across five audio conditions: clean single speaker, dialogue, crosstalk, non-native accents, and background noise. See our open-source benchmark

Speaker diarization

Percentage of audio time where the provider correctly identifies who is speaking. Assessed from vendor documentation and, where the tool is in it, our transcription benchmark.

Integration depth

We document which platforms, CRMs, and project management tools each product integrates with, from vendor docs and changelogs, and we note where users report that the data does not flow correctly on the other side.

Pricing verification

Every price carries the date we last read the vendor's page. We re-check when a reader or a vendor flags a change, and we log every correction on the corrections page.

User experience

Setup time, onboarding flow, daily workflow friction, and how the tool behaves when things go wrong (network drops, large meetings, edge cases). We note bot intrusiveness for tools that join meetings as a visible participant.

How rankings work

Rankings within each category use a weighted score across four dimensions:

35%
Accuracy / core function
25%
Pricing / value
25%
Integrations
15%
User experience

Weights shift by category. For transcription APIs, accuracy gets 45% and UX drops to 5%. For scheduling tools, UX gets 30%. We publish the weights used on each category page.

How comparisons work

Head-to-head comparisons apply the same evaluation criteria to both tools, from the same kinds of public sources. We report strengths and weaknesses for both sides. There is always a verdict, but we don't declare "winners" because the best tool depends on your specific needs.

Limitations and caveats

  • We test with English audio only unless stated otherwise. Multilingual performance may differ.
  • We use default settings for all providers. Custom vocabulary, fine-tuning, and language hints can improve results.
  • Enterprise features that require custom contracts are noted but not always tested hands-on.
  • Our test environment (audio quality, meeting size, use cases) may not match yours. Use our data as a starting point, not the final answer.
  • Tools update constantly. We re-check when a change is flagged, so a review can be out of date until it is corrected.

Independence

Some links on this site are affiliate links (always labeled). Affiliate relationships never influence rankings, scores, or editorial judgment. No vendor approves our content before it runs. A named interview subject may check a draft for factual accuracy and off-record boundaries; the piece says when that happened and editorial decisions stay with us. If a vendor offers us early access to a feature for testing, we accept it but disclose it in the review.

If we get something wrong, email hello@meetingstack.io. We correct errors openly and note the correction date.

Open-source benchmarks

Our transcription accuracy benchmark is open source. You can run the same tests yourself, add providers, or contribute audio samples.