AI Running
§2 Section 2 of 6 3,088 words · 14 min

Building Your Own Plan With ChatGPT and Claude

A large language model will write you a 16-week marathon plan in about nine seconds. It will look like a training plan: numbered weeks, Tuesday intervals, a Sunday long run, a taper. It will have paces to one second per kilometre. What it won’t have is any idea whether the ramp from your current 46 km a week to its week-9 peak is survivable, because you probably didn’t tell it what your current week looks like, and because nothing in the model’s objective function punishes it for giving you a stress fracture.

That gap is the whole game. Writing plausible training plans is easy for these tools. Writing a plan that progresses you safely requires you to supply the constraints and then check the output against them. This page is about both halves: the prompt, and the audit.

Why most AI plans fail before they are generated

Ask for “a 16-week marathon plan for a sub-3:40 finish” and the model has one usable number: 3:40. So it reverse-engineers everything from the goal. Paces come from 5:13/km marathon pace. Volume comes from whatever a sub-3:40 runner typically runs, which in the training literature it has absorbed means Pfitzinger’s 18/70 or Hansons, both of which assume a base you may not have.

Here is the actual input set that changes the output. Every one of these is available to you in Strava in under two minutes:

  • Last 8 weeks of volume, week by week. Not the average. The shape matters: 38, 41, 44, 30, 48, 51, 46, 52 is a different runner from 46 × 8.
  • Longest single run in the last 12 weeks, with its duration, not just distance.
  • A recent hard effort with pace and average HR. A parkrun, a 10K, a 20-minute time trial. This is the only honest anchor for your training paces.
  • Days you can run, and which day can be long. “Five days” is useless. “Tue/Wed/Thu/Sat/Sun, long run Sunday, Saturday capped at 45 minutes” is a constraint the model can respect.
  • Injury history with dates. “Right posterior tibial tendinopathy, flared up March 2026 after three weeks of hill reps” is actionable. “Prone to injury” is not.
  • Non-negotiables in the calendar. A week in Spain in week 7, a wedding in week 11.

Supply that and you get a plan built around you. Supply the goal only and you get a plan built around the goal, which is a different and much worse thing.

A prompt that produces an auditable plan

The trick is to give the model the arithmetic rules you intend to check it against. If it knows you will be auditing the volume curve, it builds a volume curve that survives the audit.

You are writing a 16-week marathon plan. Race: 21 Feb 2026, flat
road course. I am 42, male, 5 years of consistent running.

CURRENT STATE
Last 8 weeks (km): 38, 41, 44, 30, 48, 51, 46, 52
Longest run last 12 weeks: 21 km in 2h04
Half marathon 4 weeks ago: 1:44:30 (4:57/km), avg HR 168, max HR 186
Resting HR 48. Days available: Tue, Wed, Thu, Sat, Sun.
Saturday capped at 45 min (family). Long run Sunday.
Injury history: right Achilles niggle Oct 2024, resolved, returns
if I do more than one hill session a week.
Travel: no running week of 8 Dec (week 7).

HARD CONSTRAINTS ON THE PLAN
1. Peak weekly volume no more than 70 km.
2. Weekly volume increases no more than 8% week on week, and every
   4th week is a down week at 70-75% of the preceding week.
3. Maximum 2 quality sessions per week. Quality volume (anything at
   threshold or faster, plus marathon-pace segments) must stay
   under 20% of weekly time.
4. At least 48 h between quality sessions. Never a quality session
   the day before or after the long run.
5. Long run capped at 3h00 regardless of distance.
6. Derive all paces from my half marathon result, not from my goal
   time. State the threshold pace you derived and show the working.
7. 3-week taper.

OUTPUT
A markdown table: Week | Dates | Mon-Sun sessions | Weekly km |
Quality km | % increase on previous week. Then a second table
listing the five training paces with the derivation for each.
Flag any week where you had to break a constraint and say why.

That prompt is deliberately long. It is also the single highest-leverage thing you can do, and it is reusable: swap the current state block and the race, keep the constraints. We keep a set of these, including versions for 5K, 10K, half, and mid-block repair, in the prompt library.

Note constraint 6. Riegel’s formula puts a 1:44:30 half at 104.5 × 2^1.06 = 218 min, so 3:38 marathon, meaning a 3:40 goal is defensible if the volume arrives. Threshold pace for a 104-minute half runner sits around 4:50/km, not the 4:45 a model will often invent by working backwards from 3:40. Five seconds per kilometre sounds trivial. Across 8 × 1000 m it is the difference between a session you can repeat on Thursday and one you can’t.

Audit one: the volume curve

Read the weekly km column on its own, ignoring the sessions. This is where AI plans break most often, and it takes 30 seconds.

A plan generated without volume constraints, for the runner above, typically looks like this:

Wk   1   2   3   4   5   6   7   8   9  10  11  12  13  14  15  16
km  55  60  66  72  68  75  82  88  80  90  95  88  80  60  45  30
%   +6  +9 +10  +9  -6 +10  +9  +7  -9 +13  +6  -7  -9 -25 -25 -33

Two problems, both fatal. Week 1 is 55 km when the runner’s last eight weeks averaged 44. That is a 25% step change on day one, before any progression has happened. And the down weeks are not down weeks: week 5 drops 6%, week 9 drops 9%, week 12 drops 7%. A 6% reduction is noise, not recovery. The peak of 95 km is 116% above the runner’s actual base.

The corrected curve, which the constrained prompt produces:

Wk   1   2   3   4   5   6   7   8   9  10  11  12  13  14  15  16
km  48  52  56  40  58  62  66  48  64  68  70  52  66  50  38  26
%   -8  +8  +8 -29 +45  +7  +6 -27 +33  +6  +3 -26 +27 -24 -24 -32

Week 1 at 48 km is a step down from the runner’s most recent week of 52, which is correct: you start a block from where you are, not from where the plan wants you to be. Every fourth week drops hard. The peak is 70 km, a 59% increase on the eight-week average, achieved over 11 weeks.

Cross-check this against your own data rather than trusting either table. In Strava, the Fitness graph (CTL, if you have a subscription and heart rate) should climb at no more than about 4 to 5 points per week over a sustained block. A jump from 45 to 62 in four weeks is 4.25/week and fine. From 45 to 70 is 6.25/week and is where people get hurt. On Garmin, watch the Load Ratio: the optimal band is roughly 0.8 to 1.5, and anything the watch flags as “high” for more than a few days in a row means the plan is ahead of you. intervals.icu shows CTL ramp rate for free and will draw the line for you.

Audit two: intensity distribution and hard-day spacing

Here is a real week from an unconstrained plan, week 10 of that 95 km build:

Mon  Rest
Tue  8 km inc 5 × 1000 m @ 4:35, 90 s jog
Wed  10 km easy
Thu  12 km inc 8 km @ 4:55
Fri  Rest
Sat  5 km parkrun
Sun  26 km long, last 10 km @ 5:13

Total 61 km. Add up the hard bits: 5 km of intervals, 8 km of tempo, 5 km of parkrun, 10 km at marathon pace. That is 28 km, or 46% of the week’s distance. Convert to time, which is what matters for intensity distribution: 28 km averaging roughly 4:55 is 138 minutes, and the remaining 33 km at 6:15 is 206 minutes. So 138 of 344 minutes, 40% of training time, spent at threshold or faster. The target is 20%.

Worse, count the quality days. Tuesday, Thursday, Saturday, Sunday. Four. The Saturday-to-Sunday pairing puts a 5 km race effort 18 hours before a 26 km run with 10 km at marathon pace, which is a session stacking pattern that works for elite runners inside a deliberate block and works for nobody else.

The check is mechanical. For every week in the plan, ask the model to output a “quality minutes / total minutes” column, then scan for anything over 22%. And read the day columns left to right looking for two consecutive letters that aren’t “easy” or “rest”.

Audit three: paces anchored to current fitness

This is the subtle one, and it is the reason plans fall apart in week 6 rather than week 1.

When a model derives paces from the goal time, every session is prescribed at fitness you do not yet have. The plan asks for 8 × 1000 m at 4:35 when your current VO2max pace is 4:42. You hit 4:35 for the first four reps because you are stubborn, fade to 4:48, finish the session having run a hard time trial instead of an interval workout, and arrive at Thursday’s tempo with legs that cannot hold 4:55. Two weeks of that and you are either injured or convinced the plan is too hard.

The fix is to force the derivation and check it. For the 1:44:30 half runner:

ZonePace (min/km)Derived from
Easy6:05 to 6:30Half pace + 70 to 95 s
Marathon5:13Goal, used sparingly and only from week 8
Threshold4:50~60-minute race pace, half pace minus 7 s
5K / VO2max4:36Estimated 23:00 5K from 1:44:30 half
Rep / 15004:155K pace minus 20 s

Then ask for a re-derivation every four weeks off actual data. If a 20-minute time trial in week 9 comes out at 4:38/km, threshold has moved to roughly 4:44 and the whole table shifts. Models are happy to do this and will not do it unprompted.

Audit four: the long run and the taper

The 30%-of-weekly-volume guideline for the long run is fine for 5K and 10K training and breaks completely for the marathon. At a 70 km peak, a 32 km long run is 46% of the week. At the runner’s starting 46 km, it would be 70%.

Use time instead. Cap the long run at 3h00 for a 3:40 marathoner, which at easy pace of 6:15/km is about 28 km, and accept that you will never run 32 km in training. That is a feature. The physiological return on minutes 150 to 190 of a long run is small and the recovery cost is large. Three long runs of 2h45 in the final six weeks beats one 3h30 death march and two weeks of limping.

Watch for the specific failure of a long run that stops progressing: plans that sit at 28 km for weeks 9 through 13 without varying the structure. Progression in the final block should come from what is inside the run, not its length. Week 9: 26 km all easy. Week 11: 28 km with the final 8 km at 5:13. Week 13: 26 km with 3 × 4 km at 5:13 inside it.

On the taper, two errors recur. Some plans cut volume by 25% a week but also cut intensity to zero, which loses sharpness for no recovery benefit. Others apply a three-week taper to a 5K plan, where 8 to 10 days is enough. A defensible marathon taper from a 70 km peak: week 14 at 50 km, week 15 at 38 km, race week 26 km, with the quality sessions shortened but not slowed. Race week should still contain something like 6 × 400 m at 4:36 on Tuesday.

Making the model check its own work

The highest-value second prompt is an adversarial one. Paste the generated plan back and ask for it to be attacked:

You are a physiotherapist who works with recreational marathon
runners. Below is a plan written for the athlete profile I gave
you. Find every point where it could cause injury or
non-completion. Specifically:
- List every week where volume rises more than 8% on the previous
  week, with the actual percentages.
- List every week with more than 2 quality sessions.
- List every instance of quality work within 24 h of another
  quality session.
- Calculate quality time as a % of total time for weeks 6, 10 and 13.
- Identify the single week most likely to break this athlete and
  say why.
Do not rewrite the plan. Just the findings.

Running this on ChatGPT with code execution enabled is meaningfully better than running it on the chat model alone, because the arithmetic gets computed rather than estimated. Models are poor at summing a 16-row table reliably in prose and good at it in Python. Say “compute this in code” and the numbers stop drifting.

ChatGPT and Claude do different jobs here

Both write plans. The useful difference is in what surrounds the plan.

ChatGPT’s advantage is computation and persistence. Export your Strava activities as CSV, upload it, and ask for weekly volume totals, the actual eight-week ramp rate, and the distribution of your runs by pace decile. It will write the pandas and give you real numbers off your real history, which beats typing “about 46 km a week” from memory. Memory also means a plan built in October is still context in December.

Claude’s advantage is the document. Build the plan as an artifact in a Project, keep your athlete profile and constraints as project knowledge, and edit the plan in place over 16 weeks rather than regenerating it in a new chat each time something changes. Its longer context window also means you can paste 12 weeks of lap-by-lap workout data without truncation, which matters for the re-derivation step.

Neither is trustworthy on arithmetic in prose, and both will produce the same structural mistakes if you let them work from a goal time. The provider is not the variable that matters. The constraint block is.

Feeding your watch data back in

A plan you never revise is a guess with a table around it. Every four weeks, take four numbers from your own data and re-prompt.

Aerobic decoupling on the long run. Split your longest recent run in half in Strava and compare pace-to-HR. A 2h10 run where the first half averaged 5:45/km at 142 bpm and the second half 5:47/km at 152 bpm has decoupled by about 7%. Under 5% means your aerobic base supports that duration and you can extend it. Over 5% means hold the duration and add easy volume instead.

Whether easy runs are actually easy. Filter your last 28 days by pace. If the prescribed easy range is 6:05 to 6:30 and half your easy runs come in at 5:50, you are running a moderate-intensity plan regardless of what the table says. This is the single most common reason a well-built plan produces a flat Fitness curve and tired legs.

Garmin HRV Status and Training Status. An “Unbalanced” or “Strained” status persisting past a down week is the watch telling you the ramp is too steep. Feed that in literally: “Garmin HRV Status has read Unbalanced for 9 of the last 14 days, baseline dropped from 68 to 59 ms.”

Completion rate. If you completed 11 of the last 16 prescribed sessions, the plan is written for a runner with more time than you have. Rewrite for four days, not five.

Then the re-prompt, which takes one paragraph: “Weeks 1 to 4 of the plan below are complete. I ran 46, 50, 54 and 41 km against the prescribed 48, 52, 56, 40. I missed the week-3 threshold session (work). Long run decoupling in week 4 was 4.1% over 2h05. Resting HR unchanged at 47. Rewrite weeks 5 to 8 only, keeping all original constraints, and tell me whether the 70 km peak is still appropriate.”

Getting the plan off the chat and onto your watch

A plan in a chat window is a plan you will stop following in week 3. Ask for it in a format your tools accept.

For intervals.icu, which imports plans as files and syncs structured workouts to Garmin, Wahoo and Zwift, ask for the plan as a CSV with one row per session:

date,name,type,description
2026-11-04,Threshold 3x8,Run,"3 x 8 min @ 4:50/km w/ 2 min jog; 15 min w/u, 10 min c/d"
2026-11-05,Easy 10k,Run,"10 km @ 6:05-6:30/km"
2026-11-07,Aerobic 45,Run,"45 min @ 6:05-6:30/km + 6 x 20 s strides"
2026-11-08,Long 22,Run,"22 km @ 6:10-6:30/km, flat to rolling"

For Garmin Connect, ask for each quality session written in Garmin’s own workout vocabulary: warm-up by time, repeat block with work and recovery steps defined by distance or time and a pace target range rather than a single pace. Garmin’s target ranges want an upper and lower bound, so “4:45 to 4:55” imports cleanly where “4:50” does not.

One practical note on the apps you may be leaving. Runna and Garmin Coach both adapt to what you actually complete, and that adaptation is genuinely useful. What they do not do is show their reasoning, which means when Runna drops your Thursday session you cannot tell whether it did so because of your recent load or because of a rule in its template. An AI-built plan trades away the automatic adaptation and hands you the reasoning instead. The trade is only worth it if you do the auditing work, which is why the constraint block and the four-weekly re-prompt are not optional extras.

Pull up your last eight weeks in Strava now and write the current-state block. It takes four minutes, and it is the part of the prompt library that nobody can write for you.

In this section

The supporting pages under this subject.