How We Test
Start with the disclaimer, because most sites bury theirs. We do not run a lab. No bench, no controlled environment, no automated eval harness, no repeated-trial statistics. Nobody here is timing token throughput across ten runs. If a claim on this site needed a lab to justify it, that claim shouldn’t be on this site. Hold us to it.
So what is left? Four things, and they’re worth more than they sound.
1. Every post ships a working artifact
Each guide and prompt on Automation Squad comes with something you can use: a prompt, a template, a workflow, a decision sheet, a settings checklist. That artifact gets built in the tool it describes and run before the post goes live.
This is a low bar that a surprising amount of AI coverage fails. A recommendation that’s never been executed tends to have a gap in it — the step where the export doesn’t include what you needed, the feature is admin-only, or the integration requires a plan two tiers up. Building the thing surfaces that. Writing about the thing doesn’t.
What this proves: the workflow completes and produces what the post says it produces. What it doesn’t prove: that it’s the best workflow, or that it’ll behave the same on your tenant with your permissions and your data governance settings.
2. Every spec is verified against the vendor’s own published page
Prices, seat minimums, limits, retention windows, compliance certifications. Each one gets checked against the vendor’s own pricing page, docs or trust centre — not against a competitor’s comparison table, and not against a number we saw in someone else’s article.
Two rules govern this.
- Vendor-published beats third-party.When a review site and the vendor disagree on a price, the vendor’s page wins and the discrepancy gets noted.
- Unknown beats plausible.Enterprise pricing that says “contact sales” is recorded as null, not estimated. A guessed number would flow straight into SquadScore normalisation and quietly corrupt a ranking — a far worse outcome than an honest gap in a table.
In the AI Assistants category — the one live today — two of the eight scored specs are our own editorial rubric rather than a vendor-published figure: Setup Effort and Security & Admin. Every table labels them as such. As new categories launch, the exact count may shift; the disclosure rule doesn’t.
3. Everything carries a date
Specs move. Pricing pages change on a Tuesday with no announcement, model versions get deprecated, free tiers get quietly clipped. Every spec table carries the date it was verified. Every article carries the date it was last reviewed. If you’re reading a table stamped four months ago, treat it as four months old — that’s exactly what the stamp is for. We’d rather show you a stale date than let you assume freshness we can’t back.
4. Criticism goes in, including for tools that rank well
A tool can top a category list and still have something wrong with it. Those go in the same post, in plain words, not softened into a “consideration.”
The test we apply: if a reader adopts this tool on our recommendation and hits a problem in month two, would they feel we hid it? If yes, it goes in the post.
What we don’t do yet
Named plainly, because the gaps matter as much as the method.
- No performance benchmarking. No latency, throughput or accuracy testing of any kind.
- No blind or side-by-side output comparison. Nothing here is a controlled quality test of one model’s writing against another’s.
- No long-run reliability tracking. We don’t monitor uptime or measure incident frequency.
- No security auditing. We report compliance certifications a vendor publishes. We don’t verify them, and a SOC 2 badge in a table means “the vendor says so,” nothing stronger.
- No enterprise-scale deployment testing. Nothing here has been rolled out across thousands of seats by us. Where a tool’s behaviour at that scale matters, the post says the question is open.
- No paid-tier coverage of every tool. Some assessments rest on the free or entry tier because that’s what was accessible.
Some of these are on the roadmap. None of them are true today, and this page gets updated when one becomes true rather than before.
Independence
Automation Squad currently runs no affiliate links, ads, or sponsorships — nobody pays us to rank or recommend anything (see how we’re funded). No brand can pay for a placement, a score, or a position on any list, today or if that ever changes — see the editorial policy for the commitments that bind us if it does.
If something here is wrong
Wrong prices and stale specs are the most common failure mode of a site like this, and readers catch them faster than we do. Email [email protected]with the page and what’s wrong — a link to the vendor’s own page settles most disputes in one exchange. The editorial policy covers how corrections get handled and how fast.
