What AI Form Analysis Can See From a Phone Video — And What It Can’t
You film yourself running on a treadmill, upload the clip, and ninety seconds later an app tells you your cadence is 168, your left knee collapses inward, and your “injury risk score” is 7.2 out of 10. Three of those four outputs are worth reading. One of them is invented.
That’s the honest summary of ai running form analysis accuracy in 2026. The 2D pose estimation underneath these tools is genuinely good at some things and structurally incapable of others, and the gap between those two categories is where runners get misled. If you’re following a Runna block or a Garmin Coach plan and you’ve started wondering whether your form is the reason your shin hurts in week 9, you need to know which numbers to trust before you act on any of them.
What’s actually happening under the hood
Almost every consumer form-analysis app runs some variant of the same pipeline: a pose estimation model (Google’s MoveNet or BlazePose via MediaPipe, or something built on OpenPose/HRNet lineage) finds 17 to 33 keypoints per frame, then a layer of heuristics turns those coordinates into metrics.
Keypoints are pixel positions of estimated joint centres. Not bones. Not joint axes. Pixel guesses at where a hip or knee probably is, based on what the model learned from labelled images. On a clean side-on video with good light, MoveNet’s keypoint error on hips and knees typically lands somewhere around 8 to 20 pixels on a 1080p frame. At a typical filming distance that’s roughly 1 to 3 cm of real-world uncertainty per joint, per frame.
That error budget is the whole story. Metrics that survive it are useful. Metrics that need precision finer than it are noise wearing a number.
The stuff it gets right
Cadence. This is the most robust output in the entire category, and it’s robust because it doesn’t depend on joint precision at all. Count the frames between successive foot contacts, multiply out. At 60 fps, one frame of error on a 0.35-second contact interval gives you about ±1.4 spm. Filmed at 240 fps on an iPhone, the error becomes trivially small.
Cadence is also the metric where your watch already agrees. If Runna’s video tool says 172 and your Forerunner logged 173 avg for that session, you have two independent methods converging, which is about as much confidence as you get in this field.
Overstride, measured as foot-to-COM horizontal distance at initial contact. Here’s the thing: you’re measuring a horizontal gap of maybe 5 to 25 cm between two keypoints in the same frame. The ankle keypoint is one of the better-localised points in most pose models, and correlated error between two points in the same frame partly cancels. So a tool reporting “your foot lands 18 cm ahead of your hip” is making a claim its inputs can actually support, within perhaps ±3 cm.
Worked example. Same runner, two clips, filmed a month apart at the same treadmill speed of 12 km/h:
March clip September clip
cadence (spm) 164 176
foot-to-hip at IC (cm) 22 14
contact time (ms) 268 241
shank angle at IC (deg) 12 fwd 4 fwd
Every one of those four numbers moved in the same direction and in a physically coherent way. Higher cadence means shorter steps, which means the foot lands closer beneath the hip, which reduces the braking angle of the shank and shortens ground contact. That’s not four independent findings. It’s one change, seen four ways, and the internal consistency is what makes it believable.
Trunk lean from vertical. A forward lean of 4° versus 10° is a difference of tens of pixels at the shoulder on a decently framed clip. Visible, measurable, repeatable. The caveat is camera placement: tilt your phone 3° and your trunk lean reads 3° off, every frame, invisibly. Tools that ask you to level the phone against a doorframe are asking for a reason.
Contact time and flight time, with an asterisk. At 240 fps each frame is 4.2 ms, so a 230 ms contact time carries maybe ±8 ms of error. Fine. At 30 fps each frame is 33 ms, and your 230 ms measurement now has an error bar of ±40 ms, which is 17% of the value. The frame rate you filmed at matters more than which app you used. A 30 fps clip should not be producing contact-time numbers to the millisecond, and when one does, the precision is fabricated by interpolation.
Vertical oscillation, loosely. Pelvis keypoint height through the cycle gives you something in the 6 to 11 cm range for most runners. Directionally fine for tracking yourself over time, not worth comparing against a friend’s number from a different app.
Our gait analysis pillar walks through the filming setup that keeps these measurements inside their useful range: side-on, 5 to 7 metres back, phone level and at hip height, 120 fps minimum, and a full 20 seconds so you get 30-plus strides rather than four.
What it fundamentally cannot do
Now the part the marketing copy skips.
Forces. Ground reaction force, loading rate, impact peak, bone stress, tendon load. A camera records position. Getting force from position requires knowing segment masses, segment inertias, and acceleration accurate enough to survive double differentiation. Differentiate noisy position data twice and the noise amplifies enormously: 2 cm of keypoint jitter at 60 fps becomes acceleration error in the tens of m/s². Research-grade labs solve this with force plates measuring at 1000 Hz, plus marker sets, plus subject-specific models. Your phone has none of that.
So when an app shows “impact force: 2.4× bodyweight,” it has not measured force. It has run a regression fit on population data, keyed off your mass, speed and a couple of kinematic proxies. Vertical loading rate in real runners varies by something like 40 to 50% between individuals at identical speed and kinematics. A model like that cannot tell you where in that spread you sit, and a number presented without an error bar wider than the number itself is misleading by construction.
Rotation, and therefore most of what people actually want to know. Hip internal rotation. Femoral adduction. Tibial torsion. Pelvic drop in the frontal plane while you’re filming in the sagittal. Foot pronation velocity. These are three-dimensional motions, and a single 2D camera collapses one dimension to nothing.
Concretely: a rear-view clip showing your right knee tracking inward could be genuine hip adduction with internal rotation. Or it could be the same knee position produced by your pelvis rotating 8° toward the camera, which shifts the knee’s apparent position sideways while the actual femur behaves normally. Two very different mechanics, one identical set of pixels. No amount of model improvement resolves that, because the information isn’t in the footage.
Pronation has the same problem in worse form. Rearfoot eversion is measured against the calcaneus, which is under skin, under a sock, inside a shoe. The app is tracking the back of your trainer. Shoe heel counters deform under load in ways that are not the same as your heel moving. Studies comparing shoe-mounted to skin-mounted markers find discrepancies of 3° to 8° on eversion, on a metric whose entire clinically discussed range is about 4° to 12°. The error is the size of the signal.
Injury risk. This one deserves its own paragraph because it’s the claim most likely to change your behaviour and least likely to be true.
Nobody has established a kinematic threshold that predicts running injury with useful accuracy. The prospective literature is genuinely mixed: some studies find modest associations between peak hip adduction and patellofemoral pain in female runners, others find nothing; the large prospective work on loading rate and tibial stress fracture has produced effect sizes small enough that they don’t discriminate individuals. Meanwhile the factor that does show up repeatedly is training load: how fast you added volume, how many weeks since your last break, whether you ran through the first three days of pain.
Which means the highest-value injury information for you isn’t in a video at all. It’s in the Strava data you already have. If your AI plan took you from 34 km/week to 52 km/week across three weeks while also adding a second quality session, that 53% jump in 21 days is a far better explanation for your shin than a 6° overstride. Look at your weekly-load chart before you look at your knee angles.
An “injury risk: 7.2/10” score is a composite of measured proxies and unvalidated weights. It has no established relationship to whether you’ll get hurt. Some apps have quietly softened this language over the past couple of years, from risk scores toward “form score” or “efficiency” framing. That’s an improvement in honesty and it should tell you something about what the original claim was worth.
A concrete decision rule
Use video for what it does. Here’s the split I’d actually apply:
| Metric | Trust it? | Use it for |
|---|---|---|
| Cadence | Yes | Cross-check vs watch; track over blocks |
| Foot-to-hip at contact | Yes, ±3 cm | Spotting overstride; measuring a cue’s effect |
| Trunk lean | Yes, if phone is level | One-off posture check |
| Contact / flight time | Only at 120+ fps | Comparing your own clips, same setup |
| Vertical oscillation | Directionally | Your own trend only |
| Knee valgus from rear view | No | Nothing quantitative |
| Pronation degrees | No | Nothing |
| Impact force / loading rate | No | Nothing |
| Injury risk score | No | Ignore entirely |
Pair that with the one thing video genuinely beats your watch at, which is answering “did the cue work?” Film a baseline. Try metronome running at 5% above your natural cadence for two weeks. Film again under identical conditions. If foot-to-hip drops from 21 cm to 15 cm and contact time drops 20 ms, the intervention landed, and you know it landed because two different measurements agreed.
Where phone video earns its keep is as a mirror, not an oracle. It shows you, in numbers you can re-measure next month, whether something you deliberately changed actually changed. Ask it whether your marathon build is going to break you and you’ll get an answer, confidently formatted, that the physics of a single lens never entitled it to give.