Codifiability Index

Generative AI operates on knowledge that has already been rendered in symbolic form. The Codifiability Index (CI) is a task-level, capability-blind measure of whether the knowledge a task requires can be represented in text, code, numbers, or recorded practice at all. It asks a question prior to exposure and capability: is the task knowledge available in a form through which AI can matter?

The field measures exposure. This measures the question before it.

The dominant approach to AI and work scores occupational exposure and treats occupations as bags of independent tasks. Its authors are careful that exposure is not impact, but the field still lacks an instrument for what governs the movement from the one to the other. Codifiability identifies whether the task knowledge is available in a form through which exposure and capability can become economically operative. It is a property of the task, not of any model, so it changes slowly and can be assigned before any usage is observed.

Capability: What can a model do under evaluation? — A property of the current model generation. Changes with every release.

Exposure: Could AI potentially affect this task? — The field's dominant measure. Partly a function of capability, so it moves with it.

Usage: Does AI actually show up in the task? — Recorded after the fact, task by task, in platform data.

02 · what the pilots found

Two pilots now sit behind the index. The first scored 100 tasks drawn from across the economy. The second scored complete task sets for 28 occupations, 591 tasks in all, with predictions and outcome measure frozen before any scoring began.

Extensive margin, pilot 1: 0.47. p < 0.001, linear probability model. 100 tasks, choice-based sample balanced on the outcome.

Extensive margin, pilot 2: 0.39. Clustered t = 3.73 across 28 occupations. 588 scored tasks, anchored codebook, sample not selected on the outcome.

Usage by CI quartile, pilot 2: 0.11 → 0.42. Share of tasks with any observed usage, by quartile. Monotone, as in pilot 1.

Bundles that mix: 20 of 28. Occupations whose task sets combine codifiable and tacit work. Of the eight that do not, seven are uniformly codifiable.

Pilot 1's result survives a prior-corrected rare-events specification (slope 1.99, p = 0.002), reported because the outcome is rare in the population even though the pilot sample was balanced on it.

The second pilot is a replication rather than an extension: a different sample, drawn on predicted codifiability rather than on observed usage, and a different version of the instrument. Base rates differ between the two samples, so the coefficients are comparable and the R-squared values are not. All results are within-sample associations, suggestive rather than definitive. Blind human annotation remains outstanding.

Four dimensions, one score

CI = SR × (1 − D̄), where D̄ = (PC + SK + CC) / 3

Each O*NET task description is scored on four conceptual dimensions of knowledge structure. Codifiability is high only when symbolic representability is high and the three tacit-knowledge barriers are low. The rubric scores knowledge requirements, never current AI performance, so the index does not mechanically encode the outcome it is meant to explain.

SR · Symbolic representability: Can the task's inputs and outputs be expressed in text, code, numbers, or standardized procedures? raises CI

PC · Physical causality dependence: Does performance require intervention in material processes? lowers CI

SK · Somatic knowledge requirement: Does performance depend on sensory, embodied, or experiential skill? lowers CI

CC · Collective coordination intensity: Is the task knowledge embedded in multi-actor coordination or shared practice? lowers CI

The rubric is now published

The pilot behind the working paper was scored by applying the four dimension definitions directly, on a continuous scale, without anchors. An anchored codebook (v1.1) has since been written: five levels per dimension, behavioural definitions, a binding-knowledge rule, and worked examples. Re-scoring the first pilot under the anchored scale leaves the index almost unchanged (correlation 0.94; per-dimension weighted kappa 0.83 to 0.92), so the two rounds are comparable. Everything scored since uses the anchored version, with the model identifier and date recorded for each run.

A note on aggregation

The arithmetic form was published as provisional, with a bundle-level study named as the way to settle it. That study has run, and the answer is that at the usage margin the weakest-link forms do not beat the arithmetic mean: R-squared of 0.22 against 0.21 for a geometric mean and 0.18 for a minimum, with the arithmetic form best in 65 percent of bootstrap resamples over occupations.

The reason is more interesting than the ranking. Observed usage accumulates task by task rather than requiring a bundle to complete. Tasks in the top codifiability band are 12 percent of the sample and carry just over half of all observed usage, and a worker can use a model on one documentation task while the rest of the job stays untouched. O-ring logic governs output, not tool adoption on separable subtasks. The weakest-link question therefore moves to output-side tests, productivity and the disappearance of tasks from vacancy text, rather than being abandoned.

Two qualifications belong with that result. The raw minimum is a poor estimator of a bundle's bottleneck: it is an extreme order statistic, and in these data it correlates -0.45 with how many tasks O*NET happens to list for an occupation. A robust lower quantile, which is what the statistics literature recommends in its place, performs better than the minimum and marginally better than the arithmetic mean, but that difference sits inside the noise at 28 occupations and was found by searching across seven variants after the fact. It is treated here as a hypothesis for a pre-registered test, not as a result. The arithmetic baseline stands, now with evidence behind it rather than convention.

Index version 1.2 · arithmetic baseline · two pilots · codebook published — · pilot 1 figures are from the working paper (SSRN, 2026); pilot 2 from the bundle study, August 2026

Where observed usage sits

Every dot is one O*NET task, placed by its Codifiability Index. The upper lane holds tasks with any observed AI usage in the Anthropic Economic Index; the lower lane holds tasks with none. Hover to read a task, select it to see how its score is built. The sample is balanced by design (50 and 50), so lane sizes describe the sample, not the economy.

This is the first pilot, kept as published. The second pilot scored complete task sets rather than scattered tasks, and its bundle-level findings are summarized in section 05.

Two versions of the same rubric

Task Anatomy shows tasks as they were scored for the published pilot, on a continuous scale (rubric v1.0). The dials below use the five anchored levels of codebook v1.1, which is the scale built for human coders. Re-scoring the pilot under the anchored scale leaves the index almost unchanged, with a correlation of 0.94 and per-dimension agreement between 0.83 and 0.92, so the two are directly comparable. Fine gradations turned out to carry no information, and no two coders could reproduce them.

Usage share by CI quartile

Share of tasks with any observed usage, by quartile of the CI distribution, as reported in the working paper. The rise is monotone; the pilot supports the monotone implication of the theory, not a threshold claim.

Exposed but not adopted

CI and capability-based exposure are related (r = 0.76) but not identical. Disagreement cases illustrate the mechanism: exposure rates them moderately affected, the CI rates them tacit-intensive, and no usage is observed. Illustrative, not independent statistical evidence.

Most jobs are mixed

The second pilot asked a different question: not whether codifiable tasks attract AI use, but whether jobs hold together as bundles. Twenty-eight occupations were selected in three groups, six expected to be uniformly codifiable, six expected to be uniformly tacit, and sixteen expected to be mixed. Complete O*NET task sets were scored, 591 tasks in all, of which three were too vaguely worded to score and were flagged rather than guessed. Every prediction was frozen and timestamped before scoring began.

20 of the 28 occupations turned out to be mixed, including five of the six chosen as uniformly tacit anchors. The eight that were not are themselves informative: seven of them are uniformly codifiable bundles, and roofing is the only uniform bundle at the tacit end. Uniformity is achievable when a whole job is representable, and rare when it is not.

Why this matters

If most occupations were uniform, the line between what AI absorbs and what it does not would fall between jobs, and the question would be which occupations are affected. Because most occupations are mixed, the line falls inside jobs, and the question becomes who draws it. That is a question about organizations rather than about technology.

The binding task, named

For each mixed occupation the pilot identifies the lowest-codifiability core task, the one that does not move regardless of what happens to the rest of the bundle. Fourteen of sixteen matched the pre-registered prediction. Examples: manual therapy for physical therapists, court representation for lawyers, medication administration and monitoring for registered nurses, cutting and welding repairs for industrial machinery mechanics, patient positioning for radiologic technologists.

What did not work, reported as it happened

The pilot also tested whether a weakest-link statistic of a bundle predicts occupation-level AI usage better than an average. It does not, and the aggregation note in section 03 explains why. The prediction was frozen in advance and it failed, which is the reason it can be reported at all.

Limits

Twenty-eight occupations, scored by a language model under the anchored codebook without blind human coding, against usage from a single platform. A blind retest of the first fifteen tasks in the scoring order returned weighted kappa 0.83 to 0.98 with no disagreement larger than one anchor, which is internal consistency rather than validation.

Try the rubric

Pick a task you know well and answer four questions about the knowledge it takes to do it well. The tool places your answers in the distribution of the 100 pilot tasks. It illustrates how the index is built. It is not a validated score for your task, and it is not a prediction about what AI will do with your work.

One primitive, four transformations

Codifiability sets the possibility frontier of the division of labour under generative AI. Complementarity and institutions determine the division actually drawn. The wedge between the two is measurable, and it is where the political economy lives. The index is the primitive; four successive structures convert it into outcomes.

01 · The frontier — What generative AI can absorb

Knowledge structure, measured by the CI, sets where the technology can land. Post-2022 that boundary is in motion: high-frequency vacancy text across model-release waves tests whether codifiability predicts the direction and speed of the movement, with the index as a continuous treatment intensity.

02 · The bundle — How jobs reorganize around it

Occupations are bundles of complementary tasks, not bags of independent ones. Codifying some tasks shifts the binding constraint to the tacit complements that remain, whose return should rise. Scoring complete task sets adjudicates whether a weakest-link statistic of the bundle tracks usage more closely than a simple average.

03 · The wedge — Who draws the line, on what terrain

The division codifiability permits and the division firms and labour markets draw need not coincide, and the gap between them is organization and power rather than technology. Measuring that wedge, against efficiency and control accounts of the same reorganization, on the terrain of employment protection, licensure, collective bargaining, and algorithmic-management rules, is the capstone question.

04 · The pipeline — Where the human complement is formed

If human work concentrates in what resists codification, the supply of those capabilities is not automatic: it is produced by formation systems, from general schooling through vocational training and apprenticeship to in-service learning. When AI absorbs the entry-level tasks novices learn on, the ladder that produces expert judgment is itself cut. This is where the skills question lives, and it is a decision taken in education and training systems rather than a technical fact.

The chain closes on itself: decisions taken in formation systems shape which knowledge exists in codified form, so the pipeline feeds back into the frontier. Status, honestly stated: two pilots support the instrument at the task and bundle levels; the frontier, the wedge, and the pipeline are what the instrument makes testable, not results in hand. The bundle layer now has a first measurement, and its central finding is that most jobs are internally mixed.

The instrument moved, and here is where

This page records a working instrument, so it changes when the evidence does. Every version is dated and nothing is quietly overwritten.

A second pilot, and a replication

The extensive-margin result was re-estimated on a new sample: 588 scored tasks across 28 occupations (591 attempted, three flagged unscorable), using the anchored codebook, with standard errors clustered by occupation. The coefficient is 0.39 against 0.47 in the first pilot, and the quartile pattern remains monotone. The second sample was not selected on observed usage, which was the first pilot's main limitation.

The aggregation question is answered, and not in the direction expected

The arithmetic baseline was published as provisional pending a bundle-level test. The test ran. Weakest-link forms did not outperform it at the usage margin, and the reason is that observed usage accumulates task by task rather than requiring a bundle to complete. The weakest-link claim moves to output-side tests rather than being dropped.

The rubric is published

The original pilot had no anchored codebook; the four published dimension definitions were applied directly on a continuous scale. Codebook v1.1 is the first explicit version, with five anchors, behavioural definitions and worked examples. The first pilot was re-scored under it as a bridge, and the index and the headline result both held.

A new finding that was not looked for

Most occupations, 20 of 28, mix codifiable and tacit work inside a single job, and the uniform bundles are almost all at the codifiable end. This was not a prediction of the framework, and it moves the political-economy argument from which jobs are affected to who draws the line inside them.

An error found in our own data and corrected

Under the original unanchored scoring, court reporting scored as highly codifiable because its output is a text transcript. The binding knowledge is real-time stenographic skill. The anchored codebook's binding-knowledge rule catches this, and the case is now kept as a worked example of the most common scoring error.

Two things that did not change

The formula is unchanged, including the collective coordination dimension. A structured probe of forty coordination-heavy tasks, scored once under the full rubric and once with symbolic representability in isolation, found the association between the two constructs surviving isolation scoring, so the two-factor structure appears to be a property of the work rather than an artifact of the rubric. A small rubric-context effect was found alongside it and is disclosed rather than corrected retrospectively.

Still outstanding, unchanged since v1.0

Blind human annotation with inter-annotator reliability, cross-platform usage outcomes, and any test of productivity rather than usage.

Index version 1.2 · August 2026 · previous version 1.0, July 2026

Data and method

Task descriptions and occupations from O*NET. Usage outcome: extensive margin (any observed usage) from the Anthropic Economic Index (Handa et al. 2025). Capability-based exposure labels from Eloundou et al. (2024). Scores produced by a language model under a fixed, capability-blind rubric; the full study adds blind human annotation with inter-annotator reliability and cross-platform outcomes. The sample is choice-based and balanced on the outcome, so slopes are informative within the sample while levels are not. All results are suggestive rather than definitive.

The Tacit Shift

An English-language newsletter on tacit knowledge, skills, and what resists measurement when work meets AI.

New issues arrive first on Substack.

Korlu (2026) · Exposed but Not Adopted · SSRN 6966378