A companion page, How predictable is the qualifying mark, event by event?, shows that the mark it takes to reach a final moves around a great deal between championships in a volatile event like javelin - so much that leaning on history looks like a risky basis for a selection standard. The obvious response is: why lean on history at all, when you could instead rank the field by how well they're actually throwing right now, shortly before the championships? This page tests that idea directly against two more conventional history-based methods, using nine editions of the men's javelin final at the European Athletics Championships (2002-2024), and measures how big the error actually was for each.
This page uses men's javelin data only. Whether the same pattern holds in other events has not been tested here, and should not be assumed.
Looking at who's throwing well this year does not solve the problem - it makes it worse. Ranking the field by recent form and taking the 12th-best produced a mark that was too high in all nine editions tested - by just under two and a half metres on average. A method that misses in the same direction nine times out of nine isn't unlucky. It's biased.
A simpler method does better. Just averaging the actual mark from the previous four editions came far closer to the truth - a little over half the error of the form-based method, with no consistent direction to its misses.
Just repeating last time's mark lands in between. Predicting that the qualifying mark would simply match whatever it was at the previous edition had a typical error of 1.9m - worse than averaging four editions, but still clearly better than the form-based method, and its misses were split almost evenly between over- and under-predicting (4 of 8 editions over).
None of the three methods is precise enough to rely on. Even the best of them is typically out by more than a metre, which is larger than the year-to-year improvement most senior throwers make. The mark actually required has itself varied by four metres between editions.
Form-based predictions miss around a third of the finalists. Taking the top twelve ranked by recent form only captured, on average, about two-thirds of the athletes who actually went on to make the final.
So any forecast-based judgement here carries real, measurable uncertainty. A published method with published error bars can be checked, argued with and improved; an unpublished one cannot.
Selection panels assessing whether an athlete can "progress through the rounds and reach the final" need an estimate of what that will actually take - a mark to aim at. That estimate is a forecast, and forecasts have error. This page tests three ways of producing that forecast against nine editions of the men's javelin final at the European Athletics Championships (2002-2024), and measures how big the error actually was.
Method A, form-based. Rank every athlete in the field that actually competed by their average of their own top five throws in the 53 weeks ending one week before the championships, and take the 12th-ranked athlete's figure as the predicted cut mark. This is the obvious approach, but it has a built-in advantage a genuine forecast would not have: it uses the field that actually turned up, which is not knowable in advance. If anything, this should make Method A look better than a real pre-season forecast could ever be.
Method B, historical average. Simply average the actual qualifying mark from the previous four editions of the championship. It requires no knowledge of the field at all - just the history of past results.
Method C, last time. Simpler still: predict that the mark will be whatever it was at the previous edition of the championship, full stop. No averaging, no field data - just one number, carried forward. It is included because it is the forecast anyone could make without doing any analysis at all, and it sets the bar the other two methods need to clear to justify the extra work.
This page uses javelin data only. Whether the same pattern holds in other events has not been tested here, and should not be assumed.
Static snapshot - see "Methodology and SQL" below for how these figures are derived and how a live version would be queried.
| Year | Venue | Actual mark to reach final | Field size | Top-5 avg prediction | Delta | % predicted top 12 who qualified | 4-prev-editions prediction | Delta | Last-time prediction | Delta |
|---|---|---|---|---|---|---|---|---|---|---|
| 2002 | München | 79.04 | 26 | 82.87 | +3.83 | 83.3% | - | - | - | - |
| 2006 | Göteborg | 79.24 | 24 | 80.76 | +1.52 | 83.3% | 79.04 | -0.20 | 79.04 | +0.20 |
| 2010 | Barcelona | 76.69 | 23 | 80.66 | +3.97 | 66.7% | 79.14 | +2.45 | 79.24 | -2.55 |
| 2012 | Helsinki | 78.89 | 29 | 80.49 | +1.60 | 58.3% | 78.32 | -0.57 | 76.69 | +2.20 |
| 2014 | Zürich | 78.22 | 32 | 81.32 | +3.10 | 58.3% | 78.47 | +0.25 | 78.89 | -0.67 |
| 2016 | Amsterdam | 80.70 | 31 | 82.77 | +2.07 | 58.3% | 78.26 | -2.44 | 78.22 | +2.48 |
| 2018 | Berlin | 79.74 | 28 | 81.76 | +2.02 | 75.0% | 78.63 | -1.12 | 80.70 | -0.96 |
| 2022 | München | 77.20 | 24 | 80.80 | +3.60 | 66.7% | 79.39 | +2.19 | 79.74 | -2.54 |
| 2024 | Roma | 80.52 | 30 | 80.66 | +0.14 | 66.7% | 78.97 | -1.56 | 77.20 | +3.32 |
Delta = prediction minus actual mark. Positive (red) means the method over-predicted the mark required; negative (blue) means it under-predicted. All three delta columns are coloured on the same scale, centred on zero, so the depth of colour is comparable between methods. 2002 has no 4-prev-editions or last-time figure since this dataset starts there.
Computed from the per-edition table above, not stored separately.
| Top-5 average | 4-previous-editions | Last time | |
|---|---|---|---|
| Editions tested | 9 | 8 | 8 |
| Mean error | +2.43m | -0.13m | +0.19m |
| Mean absolute error | 2.43m | 1.35m | 1.87m |
| RMSE | 2.71m | 1.62m | 2.13m |
| Over-predictions | 9 of 9 | 3 of 8 | 4 of 8 |
| Largest error | +3.97m | +2.45m | +3.32m |
Actual mark required across 9 editions: mean 78.92m, SD 1.37m, range 76.69m to 80.70m. Across the same editions, an average of 68.5% of the twelve athletes predicted by the form-based method actually went on to qualify for the final.
Solid black line is the actual mark required. Dashed red is the top-5 form-based prediction. Dashed blue is the 4-previous-editions prediction, with a shaded band of plus or minus the 4-prev method's own RMSE (1.62m) around it, to show the forecast uncertainty rather than a single deceptively precise line. Dashed green is the last-time prediction - simply the previous edition's actual mark, carried forward one edition at a time.
The 4-previous-editions method is the most accurate of the three, and close to unbiased. The form-based method over-predicted the mark required in every single edition tested - nine out of nine - by 2.43m on average. A method that errs in the same direction nine times out of nine does not have bad luck. It has a systematic bias.
The last-time method - simply repeating the previous edition's mark - sits in between: a typical error of 1.87m, worse than averaging four editions but clearly better than the form-based method, and its 4 of 8 over-predictions show none of the form-based method's one-way bias. Averaging more history evidently helps, but even the crudest possible historical guess already beats a method that uses live form data.
The mark actually required has varied by more than four metres between editions (4.01m, from 76.69m to 80.70m). Even the best of the three methods has a typical error of roughly 1.3m, which is larger than the year-to-year improvement most senior throwers achieve.
Across the nine editions, on average only 68.5% of the twelve athletes predicted by form actually made the final. Roughly one finalist in three came from outside the pre-championship top twelve. In 2012, 2014 and 2016 that was five of twelve.
Any assessment of whether an athlete will reach a final rests on a forecast with an error of at least 1.3m. A published method with published error bars can be checked, argued with and improved. An unpublished one cannot.
This page is a static snapshot rather than a live query. "Actual mark to reach final" and "Field size" are directly reproducible from the results tables this site already loads, using the same schema and conventions as the other analysis pages. The three prediction columns require a multi-step computation (ranking the actual field by trailing form, or averaging prior editions) that is documented below but has not yet been wired into a single live query on this page.
Actual mark to reach final / Field size - the smallest qualifying-round mark among athletes who went on to compete in the final, and the number of distinct athletes in the qualifying round, for each European Championships men's javelin competition:
WITH qualifiers AS (
SELECT cr.competitionId, cr.aaId, MAX(cr.mark_distance) AS qualMark
FROM athletics_wa_competition_results cr
WHERE cr.eventId IN (10229636, 10229533)
AND cr.gender = 'M'
AND cr.race IN ('Qualification - Group', 'Round 1 - Heat')
GROUP BY cr.competitionId, cr.aaId
),
finalists AS (
SELECT DISTINCT competitionId, aaId
FROM athletics_wa_competition_results
WHERE eventId IN (10229636, 10229533)
AND gender = 'M'
AND race = 'Final'
)
SELECT
YEAR(c.startDate) AS year,
TRIM(SUBSTRING_INDEX(c.venue, ',', -1)) AS venue,
COUNT(DISTINCT q.aaId) AS fieldSize,
MIN(CASE WHEN f.aaId IS NOT NULL THEN q.qualMark END) AS actualMarkToReachFinal
FROM qualifiers q
JOIN athletics_wa_competitions c ON c.competitionId = q.competitionId
LEFT JOIN finalists f
ON f.competitionId = q.competitionId AND f.aaId = q.aaId
WHERE c.type = 'E'
GROUP BY c.competitionId, year, venue
ORDER BY year;
Qualifying-round race labels are not perfectly consistent across two decades of source data, so this query has not been re-run and checked edition by edition against the snapshot table above - treat it as the documented method, not a guarantee the two will match exactly.
Method A, top-5 form average - for each competition, take every throw by every athlete who competed there, in the 53 weeks ending one week before the championship started; average each athlete's own best five throws in that window; rank the field by that average; the 12th-ranked athlete's average is the predicted mark:
WITH field_form AS (
SELECT
c.competitionId,
r.aaId,
r.mark_distance,
ROW_NUMBER() OVER (
PARTITION BY c.competitionId, r.aaId
ORDER BY r.mark_distance DESC
) AS throwRank
FROM athletics_wa_javelin_results r
JOIN athletics_wa_competitions c ON c.type = 'E'
JOIN athletics_wa_competition_results cr
ON cr.competitionId = c.competitionId AND cr.aaId = r.aaId
WHERE r.eventId IN (10229636, 10229533)
AND r.date BETWEEN DATE_SUB(c.startDate, INTERVAL 54 WEEK)
AND DATE_SUB(c.startDate, INTERVAL 1 WEEK)
AND r.mark_distance > 0
),
top5_avg AS (
SELECT competitionId, aaId, AVG(mark_distance) AS top5avg
FROM field_form
WHERE throwRank <= 5
GROUP BY competitionId, aaId
),
ranked_field AS (
SELECT competitionId, aaId, top5avg,
ROW_NUMBER() OVER (PARTITION BY competitionId ORDER BY top5avg DESC) AS fieldRank
FROM top5_avg
)
SELECT competitionId, top5avg AS predictedMark
FROM ranked_field
WHERE fieldRank = 12;
"% predicted top 12 who qualified" compares the 12 athletes in ranked_field
with fieldRank <= 12 against the athletes who actually appear in that
competition's Final race, and reports the overlap as a percentage of 12.
Method B, 4-previous-editions average - simply the mean of the actual mark required (from the first query above) at the four most recent prior editions of the championship:
SELECT
e.year,
AVG(prev.actualMark) AS prev4Prediction
FROM editions e
JOIN editions prev ON prev.year < e.year
GROUP BY e.year
HAVING COUNT(prev.year) >= 4
ORDER BY e.year;
-- (in practice: take the four prior editions closest in time to e.year)
Method C, last time - the actual mark required (from the first query above) at the single most recent prior edition of the championship, with no averaging:
SELECT
e.year,
prev.actualMark AS prevEdPrediction
FROM editions e
JOIN editions prev
ON prev.year = (
SELECT MAX(p2.year) FROM editions p2 WHERE p2.year < e.year
)
ORDER BY e.year;