What the IPI measures, where the numbers come from, and the exact formulas behind them.
The Intelligence Price Index (IPI) measures the revealed price of knowledge work that generative AI is increasingly able to perform. It is a matched-model price index assembled from the posted prices of freelance "gigs" on Fiverr, covering tasks such as logo design, copywriting, software work, voiceover, translation, and video editing.
The question behind it is what happens to the market price of a cognitive task as machines become better at doing it. Instead of asking experts to score exposure or forecast disruption, the IPI reads prices that sellers themselves set and adjust inside a competitive marketplace. Its construction follows the logic of the Consumer Price Index (CPI), with the basket redefined from household goods to units of cognitive labor.
Two audiences make use of the index. Freelancers who sell standardized services, together with the buyers who commission them, can see where the going rate for a given kind of task has moved. Researchers studying how AI reshapes the labor market gain a task-level price series that conventional statistics rarely provide.
All prices come from a single marketplace (Fiverr, described below), though the offerings themselves are standard across both Fiverr and Upwork.
Every series is set to 100 index points in the base quarter, 2020 Q1, so a
level records the cumulative percentage change in a category's posted prices since then rather than a
dollar amount. A reading of 130 index points places current list prices 30% above their 2020 level, and
the Δ'20–'26 column reports that figure for the latest quarter, largest movers first. The
chart opens on a single category, since a composite of one line is uninformative; selecting further
categories rebuilds the weighted composite and draws it alongside them.
What the levels track is the price of work that generative AI performs with increasing competence. In this pilot every category has risen even after adjusting for inflation, but the spread is wide — from design at +23.2% over six years to audio at +154%, with the composite at +40.7%. An AI effect would show up in the ordering of those categories rather than in the height of any one line — a category rising more slowly than its neighbours is where substitution is most plausible.
But this pilot cannot support that ordering, and we do not claim it. We hold each category to a stated precision standard — within ±5% at 95% confidence in the latest quarter — and six of the seven miss it:
| Category | Real level, 2026Q1 | ±95% | Meets ±5%? |
|---|---|---|---|
| Translation | 236.3 | ±29.2% | no |
| Coding | 197.3 | ±17.1% | no |
| Audio | 254.2 | ±13.9% | no |
| Video | 209.5 | ±11.9% | no |
| Writing | 159.2 | ±8.3% | no |
| Marketing | 232.2 | ±7.7% | no |
| Design | 123.2 | ±4.8% | yes |
| Composite | 140.7 | ±3.7% | yes |
The composite passes while six of its seven components fail, which is not a contradiction: it is a review-weighted average and design carries about 71% of the weight, so the basket inherits design’s precision. The headline +40.7% is therefore on firmer ground than any individual category line except design’s.
The consequence is concrete: the three categories at the top of the change column — audio (254.2 ±13.9%), translation (236.3 ±29.2%) and marketing (232.2 ±7.7%) — have intervals that overlap one another completely, so which of them is “highest” is not something these data determine. Only design is far enough from the rest, and precise enough, to be called apart from the pack. Sorting the table by Δ produces a ranking; that ranking is mostly sampling noise, and the ±95% column is there so you can see how much.
Two further reasons not to read the ordering as an AI signal. First, it is not what an exposure story predicts anyway: translation and audio are among the tasks language models handle most directly, yet they sit at the top of the change column. Second, price is only one margin. We separately measured whether sales rates or gig dormancy broke after ChatGPT’s release, and found no detectable change in any category — but with bounds so wide (±23% to ±66%) that the test settles nothing either way. Suppressing the six imprecise categories would leave a one-category site, so we publish them all with their bands showing and the standard stated. Reaching ±5% needs roughly 900 matched gigs for writing, 1,100 for design and 1,600 for video — and about 7,400 for coding, whose prices carry less information per gig — against the low hundreds collected here; that is a target for the full-scale collection, not something this pilot achieves, and one that has to be sized on the worst category rather than the average. The figures default to real (inflation-adjusted) terms and can be switched to nominal on the chart, so consult inflation and limitations before drawing a conclusion from a single number.
The series runs quarterly from 2020 Q1 to 2026 Q1, twenty-five quarters in all,
indexed to 100 at the start. Archived prices reach back to roughly 2011, yet the chart
begins in 2020 for two reasons. Coverage before that year is sparse and irregular, with too few
matched gigs per quarter to trace a stable line. Opening in 2020 also fixes the baseline before
generative AI came into wide use, since ChatGPT launched in November 2022, while keeping the
sample dense enough to trust. Quarterly buckets replace months because each quarter pools more
matched gigs, which steadies the series against the noise a monthly frequency would introduce.
Earlier history may be added as coverage improves.
Because the window straddles November 2022, the quarters on either side of that date are the natural place to look for a substitution signal in exposed categories such as writing and design. Two confounders complicate that reading. The opening quarters coincide with the COVID shock of 2020 and 2021, which moved freelance demand for reasons unrelated to AI, so the earliest movements should not be taken as an AI story. General inflation also presses on posted prices across the whole window — the real series now removes that particular confounder, and it turns out to account for about half of the nominal rise — but the platform's own growth, its shift in seller mix, and sellers repricing as they accumulate reputation all remain. The index records how prices moved. It does not, on its own, establish why. Disentangling an AI effect from these other forces is the task of the accompanying paper, and the relevant limitations appear below.
The unit of observation is a gig, meaning a single well-defined task offered at a posted price, for instance "design a minimalist logo" or "translate 500 words from English to Spanish." Prices include the Basic tier, the Standard tier, and the Premium tier. Because a gig denotes a standardized task, its price can be followed across periods much as a CPI item is. Each observation, finally, counts only against the same gig's own earlier price, which isolates genuine price movement from shifts in the mix of sellers.
Seven categories are tracked at present: design, writing, marketing, coding, video, audio, and translation. Each individual freelancer's subcategory is also graphed so its own price change can be observed over time. A category qualifies on two grounds. Its work must recur as a standardized posted-price gig across many sellers, so that individual gigs can be matched to their own histories. It must also be plausibly exposed to generative AI, meaning something these models can increasingly do or assist with. Each gig is assigned to a category from its archived page. Work that is largely manual, performed in person, or not cleanly packaged as a fixed-price gig, such as data entry, virtual assistance, or general consulting, falls outside the index, either because it resists standardized pricing or because its exposure to AI differs in kind. The set is expected to expand and makes no claim that these seven exhaust the AI-exposed portion of knowledge work.
Every price shown is an actual historical Fiverr list price, recovered from the Internet Archive's Wayback Machine, the public web archive that has snapshotted Fiverr gig pages since roughly 2011. Nothing is surveyed and nothing is imputed. The figures are simply those sellers posted, as preserved in archived copies of their pages, aligned over time.
From 60 million archived URLs to a clean price panel
Reaching a usable series from the raw archive takes several stages, each narrowing a large pile of entries toward gigs whose prices can be followed reliably:
| Stage | What happens | Result |
|---|---|---|
| CDX retrieval | Query the Wayback Machine's index for every archived
fiverr.com gig URL. | 60M entries |
| Dedup & classify | Collapse to unique (URL, month) snapshots; tag each gig with a service category. | 22.7M unique |
| Longitudinal filter | Keep sellers with enough history (≥5 monthly snapshots spanning ≥2 years). | 48,643 sellers |
| Stratified sample | Draw a representative pilot sample of sellers and list their snapshots to download. | 500 sellers · 26,603 snapshots |
| Download | Fetch the archived HTML from the Wayback Machine (rate-limited, with retries). | 22,632 pages (85%) |
| Price extraction | Parse the Basic price out of each page (see below). | prices 2011 to 2026 |
| Matched panel | Keep only gigs seen in two or more periods, so each price change is measured against the same gig's own past. | matched gigs |
How a price is read off each page
Fiverr redesigned its pages repeatedly over the years, so extraction proceeds through a cascade of four methods, attempting the most reliable first and falling back as needed:
| Method | Era | How the price is found | Share |
|---|---|---|---|
packageList JSON | 2020+ | Embedded JSON array, price in cents | 72.9% |
| Old-style JSON | pre-2017 | JSON with price as a dollar string | 15.2% |
| Dollar fallback | all eras | $X pattern in the page text | 11.2% |
HTML <span> | 2018 to 2020 | class="price" DOM element | 0.7% |
Shares are of the 22,632 pilot pages above. The recent crawl, added later,
is almost entirely packageList.
Pages that look like gigs but are not
A Fiverr gig lives at fiverr.com/<seller>/<slug>, so the crawl keys every
observation on that two-segment path. Some of Fiverr's own section pages share the shape —
/hire/<category> (Pro category directories) and /agencies/<name>
— where the first segment is a reserved site section rather than a seller handle. Those pages
carry no packageList, so extraction fell through to the dollar fallback and picked up the
page's budget-filter default, which is not a price at all. Fiverr changed that default from
$1000 to $500 between 2024Q4 and 2025Q1, which injected a spurious
−50% step into every affected category. All such rows are now excluded at panel
construction — 3,846 of 37,782 observations, entirely within the recent crawl. The exclusion keys
on the reserved path segment rather than on the extraction method, because the dollar fallback also
recovers genuine prices from pre-2017 pages, where it is the only method that works.
What this site shows specifically
The full study spans 2011 to 2026. The figures on this page draw on the full-history
quarterly build, in which the pipeline output is aggregated into quarters and charted from
2020 Q1 to 2026 Q1, a span of twenty-five quarters. The Gigs column
on the index page reports how many matched gigs sit behind each category.
Because the sample is pilot-scale, the limitations below should temper any
strong reading of an individual category.
Three alternatives suggest themselves, and each falls short for this purpose. Surveys record what respondents believe prices are doing, a signal that recall, sentiment, and the composition of those who answer can all colour. Posted prices carry none of that mediation. Official wage statistics, such as the BLS series, are genuine measurements, but they arrive aggregated and lagged and do not resolve to the specific AI-exposed gigs of interest. One cannot locate the price of "a minimalist logo" or "500 words from English to Spanish" within them. Fiverr's packaged, posted prices supply exactly that: a task-level list price that can be matched to its own past, quarter after quarter.
The IPI is a matched-model index, following the approach the BLS applies to CPI items that are difficult to quality-adjust. The series actually plotted on the home page is a GEKS-Jevons index, deflated by CPI-U — steps 5 and 6 below.
Steps 1–4 come first anyway, because they are the clearest way to see what a matched-model index does: compare each gig only with its own past, average those changes geometrically, weight the categories together. They build the index as a chain, which is the textbook construction and also, on data as irregularly sampled as an archive, the one that fails. Step 5 explains that failure and replaces the chain; step 6 converts the result from dollars to constant purchasing power. The chained series of steps 1–4 is not plotted anywhere on this site — it is retained in the underlying data only as the comparison that motivates step 5. The composite is recomputed live in the browser whenever the basket changes.
Step 1 · Price relatives (same gig, period to period)
For every gig i seen in two consecutive quarters, take the ratio of its later price to its earlier one:
pi,t is the price of gig i in quarter
t, taken as the median Basic price when a gig has several snapshots in that quarter. Ratios
falling outside the band 0.1 to 10× are treated as data errors and discarded, and a
category-quarter enters only once at least 3 gigs are matched.
Step 2 · Category index (Jevons, chained — the comparison series)
Within a category, that quarter's price relatives are combined through a Jevons index, the geometric mean of the relatives, and chained onto the previous quarter's level:
Ict is the price index for
category c in quarter t. Sc,t denotes the set of
gigs in category c matched between t−1 and t, and
|Sc,t| its size. Multiplying the relatives and raising the product
to the power 1/n returns their geometric mean. The base quarter (2020 Q1) is fixed at
100 index points.
Step 3 · Composite IPI (weighted geometric mean)
The category indices are combined into a single figure by a Törnqvist-style weighted geometric mean:
wc is the weight of category c (see below). Averaging the log indices and exponentiating is what makes the composite a geometric rather than an arithmetic mean. Only the categories currently selected enter the sum, which is why the composite responds as the basket is toggled.
Step 4 · The reported change
The change shown for a category is the percentage difference in its index from the base quarter (2020 Q1, period 0) to the latest quarter (T):
The same expression yields the composite's Δ'20–'26 figure when the
composite index is substituted for a category's index.
Step 5 · Correcting for chain drift (the GEKS-Jevons index)
Steps 1–4 build the index as a chain: 2020Q1 is compared with 2020Q2, that with 2020Q3, and so on, each link multiplied onto the last. Because Wayback snapshots are irregular — some gigs are captured often, others only once every year or two — the set of gigs available for each link is different, and the small errors that introduces do not cancel. They compound along the chain, a well-documented problem called chain drift: by 2026 the chained index reads roughly twice what the drift-free estimate does, even though both use exactly the same prices.
The chart on the home page therefore does not use the chain at all. It shows a GEKS-Jevons index instead. The intuition: rather than walking from 2020 to 2026 one quarter at a time, compare every pair of quarters directly, then average all the routes between them. To get 2020Q1 → 2026Q1 you can go direct, or via 2022Q3, or via 2024Q1, and so on — each route uses a different set of gigs, so no single one is privileged. Averaging over all of them yields one consistent set of levels with no chain left to drift along. Formally, the direct comparison between two quarters s and t is the average log price change over the Ns,t gigs seen in both (a Jevons index):
and the index for quarter t against base quarter 0 is the geometric average of the direct route and every indirect route through a link quarter ℓ:
Each gig’s own price level cancels inside the ratio
pi,t/pi,s, so a $5 gig and a $500 gig can be pooled — only their
percentage moves count. A quarter pair is used only where at least 3 gigs are seen in both; a quarter
appears in the chart only if it can be reached from the base period this way, and we leave it out rather
than guess otherwise. Confidence bands come from a bootstrap that resamples gigs, so they widen where a
quarter rests on fewer or thinner comparisons (translation and audio most of all). Unlike a regression-based
correction, GEKS never imputes a price for a gig-quarter it did not observe — it uses matched pairs
only.
Step 6 · Deflating to real terms (the default view)
Everything above yields a price in dollars, and the dollar itself changed over this window. To ask whether intelligence work genuinely became more expensive, rather than whether the currency became worth less, each quarter's level is divided by the change in consumer prices since the base quarter:
CPIt is the quarterly average of CPI-U — the US
Consumer Price Index for All Urban Consumers, city average, all items, seasonally adjusted, published
by the Bureau of Labor Statistics (FRED series CPIAUCSL) — and CPI0
its value in the base quarter, 2020 Q1. The result is stated in constant 2020 Q1 dollars.
The seasonally adjusted series is used because the index is read quarter over quarter, and adjustment
removes the seasonal wave that would otherwise appear in every real quarterly change; the unadjusted
series was checked as a robustness test and never diverges by more than 0.36%.
Because the deflator is a single constant per quarter and carries no sampling error of its own, the
bootstrap confidence bands are identical on the real and nominal series — deflation moves
the level, not the precision. See question 9 for what this changes and what
it does not.
Yes — the default view is inflation-adjusted. The chart opens on the Real series, which divides the index by CPI-U (the US Consumer Price Index for All Urban Consumers, city average, all items, seasonally adjusted, published by the Bureau of Labor Statistics). A real reading answers “how much of an ordinary basket of goods does this gig cost?” rather than “how many dollars does it cost?”. The Nominal toggle shows the undeflated dollar series, with CPI-U drawn alongside it as a dashed grey line so the gap between the two is visible directly.
The adjustment matters a great deal here. US consumer prices rose 26.8% between 2020Q1 and 2026Q1, and the composite IPI rose 78.4% in nominal terms over the same window — so roughly half of the nominal increase is the dollar losing value, not intelligence work becoming more expensive. In real terms the composite rise is +40.7%. Some categories change character substantially once deflated: design is +56.1% nominal but +23.2% real over six years, and writing goes from +101.8% to +59.2%.
Two caveats worth stating. First, deflation removes general inflation and nothing else — it is not a control for sellers raising prices as they accumulate reviews and reputation, for shifts in Fiverr's seller mix, or for the fact that a matched-model index follows surviving gigs rather than new entrants. Those are separate corrections. Second, the BLS did not publish a CPI-U figure for October 2025, so the 2025Q4 deflator interpolates that one month from its neighbours; the quarter is flagged in the underlying data and in the chart tooltip.
What speaks most directly to substitution by AI is a fall in gig prices, or a rise slower than general prices, against an inflationary backdrop — which is what the real series is built to show. Note that the featured gig price histories elsewhere on the site are the nominal dollar amounts actually posted on Fiverr and are not deflated.
Weights are intended, in the manner of CPI expenditure weights, to reflect how much economic activity each category carries. Transaction volume is proxied by review counts, on the premise that a gig accumulates reviews roughly in proportion to its sales. For each category the maximum observed review count of its gigs is summed, and the totals are normalized to add to 1:
Rc is the total review volume of category
c. In the present sample design carries most of the basket, near 71%, with writing
next at about 11% and marketing, coding, video, audio, and translation dividing the remainder. Each
category's weight appears in the Weight column on the index page.
The geometric mean is symmetric under reversal: a price that doubles and then halves returns to its starting point, where an arithmetic mean of the relatives would record a spurious net rise. The Bureau of Labor Statistics likewise uses it for elementary CPI aggregates, which keeps the IPI comparable with how headline inflation is actually measured.
That same construction limits how far a handful of sellers can move the index. Since each gig is
compared only with its own past, sellers entering or leaving the sample cannot by themselves shift a
level. The geometric mean then damps extreme ratios far more heavily than an arithmetic mean would,
and two guardrails discard relatives outside 0.1 to 10× and require at least 3
matched gigs before a category-quarter counts. No individual seller's price change moves a
category by much, and averaging across categories dilutes it further. The more serious threat is
thin coverage, too few matched pairs in a given cell, rather than manipulation by any one
participant, and it is flagged in the limitations below.
The composite is rebuilt in the browser from the category indices and their weights each time the basket changes, applying the Step 3 formula above and renormalizing over whatever categories are selected. This makes it possible to ask what the index looks like for, say, design and writing alone, without relying on a server to recompute it. The checkboxes toggle categories, and the All / None links select or clear the set at once.
The interface also opens up the material underneath the composite. Each category can be expanded to its leading freelancers, and an individual gig can be opened to show its own posted price across time. Because a matched gig is nothing more than one seller's price set against its earlier self, this per-gig view is the most direct check on what the index summarizes. It also shows plainly why thinly covered categories look flat: when few gigs are archived in a quarter, there are simply not many lines that can move.
+78.4% nominal against
+282.9% chained). Different drift-free estimators still disagree with one another by
several index points in the thinner categories, so the direction of movement and the
ordering across categories carry more weight than any absolute magnitude.100 index points for long stretches, which reflects missing matches rather than genuine
price stability. The earliest quarters and the smallest categories (translation, audio) are most exposed
to this.The series can be regenerated from the project's analysis pipeline under code/. Three
steps produce what this site shows: 21-geks-index.py estimates the drift-free GEKS-Jevons
indices and their bootstrap standard errors, 23-real-index.py fetches CPI-U and deflates
them to real terms, and 18-build-site-data-long.py splices the historical and recent
segments, rebases them, and writes docs/data.json. The page loads that file and recomputes
the composite on the client. README.md and GUIDE.md alongside this page
document the data contract and the build steps, though both still describe an earlier revision of that
contract; docs/data.json itself is the authoritative version.
The GEKS implementation was checked against an independent reference package
(PriceIndexCalc) on the same panel and reproduces it exactly, to within
0.0000 index points, on every category that package is able to process.
Updates come from rebuilding against fresh Wayback Machine snapshots rather than from a live feed, so the index moves when the pipeline is re-run, not continuously. Each build stamps the page with a generation date and the quarters it covers, so the currency of the displayed series stays visible. The most recent quarter depends on pages that have actually been archived and matched, which means it can shift a little as further snapshots arrive and settles as that quarter fills in.
The project is open. Code, data-build scripts, and this page live at github.com/AISmithLab/IntelligencePriceIndex. A misread price, a misclassified gig, or a bug in the pipeline can be reported as an issue or a pull request there. Suggestions on methodology are equally welcome, since the index is meant to be auditable, and corrections from people who price this kind of work on Fiverr and Upwork improve it.