Every term from the board, one plain sentence each

Plain-English glossary

If a word on the board or in the paper sounded like another field, it is here. Most of the machine-learning vocabulary comes down to one idea: when you are allowed to look at the data. The statistics are the ordinary kind. Dip in where you need to; nobody reads this straight through.


Each card has the same shape: the term as we used it, a plain sentence, why it mattered here, and one or two things to read if you want more. Citations marked with a local PDF are in the project corpus. Others are standard sources used in the manuscript; visualization books in the design note are flagged separately because those PDFs are not in the corpus.

Where to start Splits and peeking Models and prompts Agreement Inference Forum layer

If you read only four things

  1. Törnberg (2024) on how to annotate text with a language model without fooling yourself. Local PDF.
  2. Gilardi, Alizadeh, and Kubli (2023) for the claim that models can beat crowd workers — and then our result that six prompts still fell short of a human ceiling set in advance. Local PDF.
  3. Koenig et al. (2023) in Personnel Psychology for ML in an I-O voice. Local PDF.
  4. Rosso, Dekas, and Wrzesniewski (2010) for the 2 × 2 itself. Local PDF.

A fifth, if you want the correction we specified and then did not use: Egami, Hinck, Stewart, and Wei (2023) on design-based supervised learning. Local PDF. We did not apply it because the Study 1 predictors are human codes, not machine labels.

Splits and peeking

The poster and the rest of this site use plain words; the manuscript uses the machine-learning ones. Same thing, two names: practice cases = training set; try-out cases, looked at once = development set; set-aside cases = sealed holdout / test set; locked instructions = frozen prompt; pass mark = acceptance bar; agreement score = positive specific agreement (PSA); a false present call = false positive on an agreed-absent cell; dated record of the run = provenance (model pin, prompt hash, git SHA).

Train, development, and test (practice, try-out, set-aside)

Imagine you lock three drawers before anyone codes. The small drawer (train, about 10%) is the sandbox: try wordings, try examples. The middle drawer (development, about 20%) is opened once, to pick among candidates whose instructions were already locked. The large drawer (test, about 70%) stays shut until the method is final. You do not improve the recipe after tasting the confirmatory meal.

That is ordinary machine-learning hygiene. We ported it to survey narratives so a language model could not be tuned on the same cases used to claim success.

Sealed holdout (the set-aside file)

The test drawer is not “we will be good.” It is a file nobody scores while the instructions are still being chosen. After no set of instructions reached the pass mark on the try-out cases, none was selected and the paper’s results were fixed on human codes. Humans are the gold standard. Only then was one descriptive look taken at 65 cases from the drawer, with the setups, statistics, and reporting rule written down and time-stamped before the first call; it is reported as a number on cases the models had never seen, not as validation: agreement on agency .81 and .80, on others-directed meaning .75 and .77 for the two top setups. The pre-set rule for scoring the other 128 cases was not met, so they stay untouched. The order matters: nothing upstream of that look could change because of it.

One-shot evaluation (one look)

You lock the instructions. You run them once on the try-out cases. You do not rewrite the instructions after seeing the score. If you do, development has become a second training set and the number is no longer a test.

Models and prompts

Large language model (LLM)

A text system trained to continue and classify language. Here it was asked to apply a codebook, not to chat. It is a coder with no memory of your theory unless you put the theory in the instructions.

Engineered prompt (a set of model instructions; what we used to call an arm)

The written instruction given to the model: the codebook, the output format, and any examples. Six sets were locked: decision-first, a revision of that, per-facet (thirteen calls per story), reason-first, decision plus evidence, and an alternate codebook. None reached the pass mark.

Human composite gold (the human answer key)

Two trained people coded each story. Disagreements went to a third, blinded pass. A facet is present if the composite says present. Dimensions are present if any of their facets are present (logical OR). With no model setup accepted, these human codes were the Study 1 predictors.

Design-based supervised learning (DSL)

If you regress an outcome on noisy machine labels, the slope is biased. Egami and colleagues describe a design that uses a small set of human labels to correct that bias. We specified the procedure in case machine labels were used. They were not. DSL was not applied. Mentioning it is a disclosure, not a method we ran.

Agreement, pass marks, and why communion was sidelined

Positive specific agreement (PSA), also F1: the agreement score

Kappa can look terrible when a code is rare, even if both people mostly say “absent.” PSA asks a narrower question: when someone says the facet is present, how often does the other person agree? That is the same arithmetic as F1 on a binary code. We used it because the coding target is categorical presence, not a correlation.

Acceptance standard / acceptance bar (the pass mark)

A set of instructions passed only if its agreement score with the human answer key reached the lower bound of the human–human 95% CI: ≥ .910 for agency, ≥ .805 for others-directed meaning. Communion (human .48) and self (human .52) had no pass mark, because you cannot validate a model against a human pair that does not agree. Humans are the gold standard: if the model does not clear those marks, the paper stays on the human codes.

In one sentence: we would only replace the humans if the model was at least as good as the worse end of what those humans already do, on the dimensions they can see.

Salience and ICC

Humans made present/absent calls. A 0–100 “how central is this?” rating was attempted with the models. Intraclass correlation (ICC) measures whether those ratings agree. Three prompts reached only an exploratory band (ICC(2,1) = .58–.62). We do not report salience as a graded Study 1 measure.

The regression layer (you already know most of this)

OLS with HC3

Ordinary least squares, with a standard-error correction that does not assume equal residual variance. Same models you would run in R or SPSS. Age, work centrality, occupation block, and the Time 1 baseline of the outcome when it existed.

One-tailed tests, two-sided intervals

Registered directional predictions were tested in the predicted direction at α = .05. The printed 95% confidence intervals are still two-sided. If the interval includes zero, a one-tailed p above .05 is not a surprise.

Equivalence tests (TOST) and Bayes factors

A nonsignificant p can mean “we did not have enough data” or “the effect is small.” Two one-sided tests (TOST) ask whether you can reject effects as large as a chosen bound (here |β| = .20 and .30). A Bayes factor BF01 is a rough odds ratio favoring the null over the focal predictor. For the three-dimension block over Time 1 self-reported meaning, BF01 = 481.5. That is the sentence “the codes add nothing next year beyond asking people how meaningful the work already felt.”

Benjamini–Hochberg screening

After the registered tests returned null, we searched about 1,600 tests. Each declared family was screened so the expected fraction of false discoveries among the survivors is 10% (q = .10). Survivors are exploratory. They are not a second confirmatory family.

Quantile regression

Ordinary regression describes the mean. Quantile regression describes a percentile. The code–WAMI association was larger at the 25th percentile of meaningfulness than at the 75th. We treat the shape as the finding, not the exact slopes. WAMI is ceiling-compressed.

The forum layer (different data, different rules)

Meaning event

A passage in a post where one of the 13 facets is enacted, frustrated, or wished for, as read by one language model in one pass. Each event also records what the writer blames or credits (a named AI system, a system rollout, an optimization practice such as ratios or charting load, or nothing), what happened to the task, the emotion, and who is held responsible. This is how we found that nurses name the apparatus, not AI. No human coded forum posts, so an event rate is a rate of model-assigned codes; the false-positive rate on forum prose is unknown.

Same post is not “before”

Posts that name the apparatus also carry quit-intent and lack-of-control flags more often. That is one coder reading one complaint: the same sentence can produce the object code and the flag. Read both directions. Fifteen percent of apparatus posts carry a quit flag; thirty percent of quit-flagged posts name an apparatus, so seventy percent do not. Neither number says what came first. “Precedes” needs the same writer, an earlier post and a later one, and real posting dates. We have not run that.

Dictionary hit rate

A closed list of words and phrases. If a post contains one, it counts as a hit. Fast, auditable, and not a Rosso facet score. We used dictionaries for care talk, mastery talk, optimization, AI/automation mentions, and leaving or strain. On a nursing board the care list mostly fires on ordinary patient talk, so a hit says “this post is about nursing,” not “communion is present.” We keep dictionaries as archive description (how often does AI vocabulary appear at all: under 1% of posts in any year) and draw no year-trend conclusion from them.

Capture year versus posted date

The Internet Archive stamps when it photographed the page, not when the nurse wrote the post. A 2024 snapshot can contain an older thread. We say “capture year” on purpose. It is not an exposure measure and it is not a panel of the same writers.

Transfer validation

A codebook built on structured survey stories does not automatically work on messy forum posts. Transfer validation means humans code a fresh sample of those posts (here, 200 already drawn) and you compare the model to those humans before you trust a year contrast. The codebook itself is the one two humans built on the stories. The transfer comparison is the unfinished step: the sheet was drawn and, by PI decision, not coded for this packet. Forum numbers here are that human-built instrument applied by one model to a new text type.

Item response theory (IRT), mentioned so you can ignore it

A homemade 2PL model was fit to unvalidated, binarized draft salience scores, and another to the dictionary hits. Neither is a meaning theta you can take to a client: the salience version has reliability .46, and the dictionary version reproduces the hit count (rank correlation .99) with reliability .20. Both were removed from the board and from this poster page; they are mentioned here only so nobody reconstructs them from the repository and calls them a finding.

A 10-minute path through the citations

You already know…Open this
Interrater reliability, not MLCicchetti and Feinstein (1990); then Törnberg (2024)
OLS and p values, suspicious of “n.s.”Lakens (2017); Wagenmakers (2007)
I-O selection / ML hypeKoenig et al. (2023)
Meaningful work, not methodsRosso et al. (2010); Bunderson and Thompson (2009); Bankins and Formosa (2023)
Vendor demo tomorrowTry this quarter and the PSA card above

Full bibliographic entries sit on the annotated readings page. Visualization sources for how the digital poster was drawn are on the design notes page and are not in the local PDF corpus.