There is a specific thing people type into Google when they've decided their form is the problem: workout form check app. Then squat form check app, deadlift form check app, gym form checker app. It's a shopping query. Somewhere between the third missed rep and the fourth tutorial, you stopped wanting to be told what a squat should look like and started wanting to know what yours looks like.
The results all promise roughly the same thing, in roughly the same words: upload a clip, get a score, get your fixes. Range of motion, tempo, alignment, control. A number out of ten and three bullet points.
Here's the problem with that output, and it's structural rather than cynical: a score and a list of cues are the two things easiest to produce without looking at your video at all. "Keep your chest up, control the eccentric, don't let the knees cave" is correct, generic advice for the squat. It applies to almost everyone almost always. An app that knows only the name of the exercise can generate it. So can one that watched every frame. The output looks identical, which means the output can't tell you which one you bought.
So don't evaluate the advice. Evaluate whether anything was measured. That's testable in about ten minutes, and here is how.
Test 1: upload the same clip twice
This is the whole thing. Take one video — the same file, not a second set — and submit it twice, an hour apart if the app caches.
You should get the same answer. Not similar in spirit: the same. If the first pass says your depth is fine and the second flags it, or the score moves from 7 to 5, then nothing in that pipeline measured your hip. Something generated an opinion about a video, twice, and the two opinions disagreed.
This is the single most common complaint in the reviews of apps in this category, and it's worth understanding why it happens rather than assuming bad faith. Measurement is deterministic — the bar was where it was, the knee travelled where it travelled, and asking twice can't change it. Generated commentary isn't. If a tool is mostly the second thing wearing the clothes of the first, running it twice is what makes that visible, and it's the one test that no amount of confident phrasing survives.
A form check that isn't repeatable isn't a form check. It's a horoscope with a rep count.
Test 2: give it a good rep and a bad rep
Film the same lift twice on purpose. One set you're genuinely happy with, one where you deliberately do the thing you know is wrong — cut the depth by a few inches, let the bar drift forward, rush the bottom.
Then compare the two outputs. They should differ, and they should differ in the direction you engineered. If both come back with the same three cues, you've learned that the cues aren't coming from your footage.
The failure mode here is subtler than test 1 and worth naming, because it's the one that flatters you: some tools grade everything gently and some grade everything harshly, and both feel like feedback. All-positive is the more dangerous of the two — a verdict of "perfect depth, excellent mobility" on a rep you know was high tells you the instrument is broken, and it tells you in the least alarming way possible. Feedback that never changes is not feedback. It's a mood.
Test 3: ask what it measured, not what to fix
Cues are opinions. Measurements are pointable. The difference matters more than it sounds.
"Keep your chest up" is an instruction whose truth you can't verify from your own video. "Your hip crease finished about two inches above your knee" is a claim you can pause the footage and check yourself with a finger on the screen. One is a coaching style. The other is a number that either matches what's on the screen or doesn't.
So ask for the second kind: how deep did I go, where was the bar over my foot, which rep changed. An app that can only answer in adjectives read the exercise name. An app that answers in positions and frames looked at the video — and, usefully, has now given you something you can catch it being wrong about.
This is also why "is my form good?" is the wrong thing to ask software, the same way it's the wrong thing to ask yourself while watching your own footage. It's a request for a verdict, and verdicts are unfalsifiable. Spatial questions have answers.
Test 4: change one thing and film again
The only test that measures the thing you actually care about, which is not accuracy — it's whether using this changes what your body does next week.
Take one cue the app gave you. Apply it on the very next set. Film that set. Submit it. Did the output move?
If yes, you have a feedback loop, which is a genuinely rare thing to own and the entire reason to bother with any of this. If no — if the second video comes back with the same score and the same three bullets despite a visibly different rep — then the loop is open, and an open loop will never make you better no matter how good the individual answers sound.
Almost nobody runs this test, because it takes two sessions instead of one upload. It's also the only one whose result you'd act on.
What none of them can do, including ours
Worth saying plainly, because a page like this is otherwise just marketing with a checklist stapled to it.
A phone films from one angle, so anything happening in a plane the camera can't see is invisible — rotation, what your left side is doing while you're filmed from the right, most of what's happening at your feet inside your shoes. Two cameras fix some of it and nobody films with two cameras.
No app can tell you whether something is going to injure you. That's a clinical judgment about a specific body with a specific history, and any tool implying otherwise is selling certainty it does not have.
And real-time correction — the overlay that shouts at you mid-rep — is mostly a demo rather than a product, for reasons we've written about at length: by the time a cue reaches you, the rep is over, and the useful version of this work happens after the set, not during it.
Also worth knowing before you conclude an app is broken: two competent human coaches will give you different answers about the same video. Disagreement between tools isn't automatically incompetence. Disagreement between one tool and itself — test 1 — is.
Run these on us
We build Flexion, so this is the part where you'd expect the checklist to quietly become a list of things we happen to do well. Instead: run all four on us. Especially test 1.
What we chose to build is the unglamorous half — you record the set, and the footage gets read frame by frame afterward, looking for the positions that actually answer a spatial question, and kept so the next one has something to sit next to. Not a coach yelling in your ear at the bottom of a squat, which we think is theater.
That's a claim, and the whole point of this post is that claims in this category are cheap. The tests are ten minutes. They cost you nothing and they're the only thing between you and a very confident number that nobody computed.
If a tool passes, the harder question is what to actually look at in the footage: what to look for in your own squat video is the list, in order.

