Where to start
Splits and peeking
Models and prompts
Agreement
Inference
Forum layer
If you read only four things
- Törnberg (2024) on how to annotate text with a language model without fooling yourself. Local PDF.
- Gilardi, Alizadeh, and Kubli (2023) for the claim that models can beat crowd workers — and then our result that six prompts still fell short of a human ceiling set in advance. Local PDF.
- Koenig et al. (2023) in Personnel Psychology for ML in an I-O voice. Local PDF.
- Rosso, Dekas, and Wrzesniewski (2010) for the 2 × 2 itself. Local PDF.
A fifth, if you want the correction we specified and then did not use: Egami, Hinck, Stewart, and Wei (2023) on design-based supervised learning. Local PDF. We did not apply it because the Study 1 predictors are human codes, not machine labels.
Splits and peeking
The poster and the rest of this site use plain words; the manuscript uses the machine-learning ones. Same thing, two names: practice cases = training set; try-out cases, looked at once = development set; set-aside cases = sealed holdout / test set; locked instructions = frozen prompt; pass mark = acceptance bar; agreement score = positive specific agreement (PSA); a false present call = false positive on an agreed-absent cell; dated record of the run = provenance (model pin, prompt hash, git SHA).
Train, development, and test (practice, try-out, set-aside)
Imagine you lock three drawers before anyone codes. The small drawer (train, about 10%) is the sandbox: try wordings, try examples. The middle drawer (development, about 20%) is opened once, to pick among candidates whose instructions were already locked. The large drawer (test, about 70%) stays shut until the method is final. You do not improve the recipe after tasting the confirmatory meal.
That is ordinary machine-learning hygiene. We ported it to survey narratives so a language model could not be tuned on the same cases used to claim success.
Read next. Koenig et al. (2023), Personnel Psychology, on prediction and leakage in personnel research. Yarkoni and Westfall (2017), Perspectives on Psychological Science, “Choosing prediction over explanation,” is the broader psychology argument. The Yarkoni paper is not in the local corpus; the Koenig paper is.
Sealed holdout (the set-aside file)
The test drawer is not “we will be good.” It is a file nobody scores while the instructions are still being chosen. After no set of instructions reached the pass mark on the try-out cases, none was selected and the paper’s results were fixed on human codes. Humans are the gold standard. Only then was one descriptive look taken at 65 cases from the drawer, with the setups, statistics, and reporting rule written down and time-stamped before the first call; it is reported as a number on cases the models had never seen, not as validation: agreement on agency .81 and .80, on others-directed meaning .75 and .77 for the two top setups. The pre-set rule for scoring the other 128 cases was not met, so they stay untouched. The order matters: nothing upstream of that look could change because of it.
Read next. Nosek, Ebersole, DeHaven, and Mellor (2018) on the preregistration revolution, Proceedings of the National Academy of Sciences. Used in the manuscript as the spirit of locking rules before numbers; not a claim that this study was a public OSF preregistration. It was an internal, git-timestamped specification.
One-shot evaluation (one look)
You lock the instructions. You run them once on the try-out cases. You do not rewrite the instructions after seeing the score. If you do, development has become a second training set and the number is no longer a test.
Read next. Törnberg (2024), “Best practices for text annotation with large language models,” Sociologica. Local PDF. Pangakis, Wolken, and Fasching (2023), Automated annotation with generative AI requires validation (arXiv:2306.00176), is the blunt version of the same warning.
Models and prompts
Large language model (LLM)
A text system trained to continue and classify language. Here it was asked to apply a codebook, not to chat. It is a coder with no memory of your theory unless you put the theory in the instructions.
Read next. Ziems et al. (2024), “Can large language models transform computational social science?” Computational Linguistics. Local PDF. Gilardi et al. (2023), Proceedings of the National Academy of Sciences. Local PDF.
Engineered prompt (a set of model instructions; what we used to call an arm)
The written instruction given to the model: the codebook, the output format, and any examples. Six sets were locked: decision-first, a revision of that, per-facet (thirteen calls per story), reason-first, decision plus evidence, and an alternate codebook. None reached the pass mark.
Read next. Same Törnberg (2024) paper. For the older few-shot idea in language models, Brown et al. (2020), “Language models are few-shot learners,” is the source everyone cites; it is a methods landmark, not a license to skip a human bar.
Human composite gold (the human answer key)
Two trained people coded each story. Disagreements went to a third, blinded pass. A facet is present if the composite says present. Dimensions are present if any of their facets are present (logical OR). With no model setup accepted, these human codes were the Study 1 predictors.
Read next. Flanagan (1954) on the critical incident technique, Psychological Bulletin. O’Connor and Joffe (2020) on intercoder reliability debates, if you want the qualitative-methods side of the same problem.
Design-based supervised learning (DSL)
If you regress an outcome on noisy machine labels, the slope is biased. Egami and colleagues describe a design that uses a small set of human labels to correct that bias. We specified the procedure in case machine labels were used. They were not. DSL was not applied. Mentioning it is a disclosure, not a method we ran.
Read next. Egami, Hinck, Stewart, and Wei (2023), NeurIPS. Local PDF. Read the abstract and the social-science examples first; the proofs can wait.
Agreement, pass marks, and why communion was sidelined
Positive specific agreement (PSA), also F1: the agreement score
Kappa can look terrible when a code is rare, even if both people mostly say “absent.” PSA asks a narrower question: when someone says the facet is present, how often does the other person agree? That is the same arithmetic as F1 on a binary code. We used it because the coding target is categorical presence, not a correlation.
Read next. Cicchetti and Feinstein (1990), “High agreement but low kappa: II,” Journal of Clinical Epidemiology. Feinstein and Cicchetti (1990) is part I. van Rijsbergen (1979), Information Retrieval, is the F1 source. Hayes and Krippendorff (2007) is the usual communication-science alternative if your reviewers ask for alpha instead.
Acceptance standard / acceptance bar (the pass mark)
A set of instructions passed only if its agreement score with the human answer key reached the lower bound of the human–human 95% CI: ≥ .910 for agency, ≥ .805 for others-directed meaning. Communion (human .48) and self (human .52) had no pass mark, because you cannot validate a model against a human pair that does not agree. Humans are the gold standard: if the model does not clear those marks, the paper stays on the human codes.
In one sentence: we would only replace the humans if the model was at least as good as the worse end of what those humans already do, on the dimensions they can see.
Read next. Törnberg (2024) again, on reporting confusion structure rather than a single accuracy number. Our practice note is simpler: one global F1 is a marketing statistic.
Salience and ICC
Humans made present/absent calls. A 0–100 “how central is this?” rating was attempted with the models. Intraclass correlation (ICC) measures whether those ratings agree. Three prompts reached only an exploratory band (ICC(2,1) = .58–.62). We do not report salience as a graded Study 1 measure.
Read next. Any standard ICC chapter will do; we used the usual Shrout and Fleiss (1979) language. The number that matters for this project is the pre-declared rule: reportable ≥ .75, exploratory ≥ .50, otherwise drop.
The regression layer (you already know most of this)
OLS with HC3
Ordinary least squares, with a standard-error correction that does not assume equal residual variance. Same models you would run in R or SPSS. Age, work centrality, occupation block, and the Time 1 baseline of the outcome when it existed.
Read next. Long and Ervin (2000), “Using heteroscedasticity consistent standard errors in the linear regression model,” The American Statistician.
One-tailed tests, two-sided intervals
Registered directional predictions were tested in the predicted direction at α = .05. The printed 95% confidence intervals are still two-sided. If the interval includes zero, a one-tailed p above .05 is not a surprise.
Read next. The confirmatory section of the manuscript methods. No special citation required beyond ordinary APA reporting.
Equivalence tests (TOST) and Bayes factors
A nonsignificant p can mean “we did not have enough data” or “the effect is small.” Two one-sided tests (TOST) ask whether you can reject effects as large as a chosen bound (here |β| = .20 and .30). A Bayes factor BF01 is a rough odds ratio favoring the null over the focal predictor. For the three-dimension block over Time 1 self-reported meaning, BF01 = 481.5. That is the sentence “the codes add nothing next year beyond asking people how meaningful the work already felt.”
Read next. Lakens (2017), “Equivalence tests,” Social Psychological and Personality Science — the most readable TOST explainer in psychology. Wagenmakers (2007), “A practical solution to the pervasive problems of p values,” Psychonomic Bulletin & Review, for the BIC approximation we used. These supplements were not preregistered; the decision rules were fixed before the numbers were computed.
Benjamini–Hochberg screening
After the registered tests returned null, we searched about 1,600 tests. Each declared family was screened so the expected fraction of false discoveries among the survivors is 10% (q = .10). Survivors are exploratory. They are not a second confirmatory family.
Read next. Benjamini and Hochberg (1995), Journal of the Royal Statistical Society: Series B. If you want the “why we still explore” essay rather than the algorithm, Tukey (1980) and Scheel, Tiokhin, Isager, and Lakens (2021).
Quantile regression
Ordinary regression describes the mean. Quantile regression describes a percentile. The code–WAMI association was larger at the 25th percentile of meaningfulness than at the 75th. We treat the shape as the finding, not the exact slopes. WAMI is ceiling-compressed.
Read next. Koenker and Bassett (1978), Econometrica, is the original. You do not need it to use the poster sentence.
The forum layer (different data, different rules)
Meaning event
A passage in a post where one of the 13 facets is enacted, frustrated, or wished for, as read by one language model in one pass. Each event also records what the writer blames or credits (a named AI system, a system rollout, an optimization practice such as ratios or charting load, or nothing), what happened to the task, the emotion, and who is held responsible. This is how we found that nurses name the apparatus, not AI. No human coded forum posts, so an event rate is a rate of model-assigned codes; the false-positive rate on forum prose is unknown.
Read next. Rosso, Dekas, and Wrzesniewski (2010) for the facets. Törnberg (2024) for why a code that works on one text type can fail on another.
Same post is not “before”
Posts that name the apparatus also carry quit-intent and lack-of-control flags more often. That is one coder reading one complaint: the same sentence can produce the object code and the flag. Read both directions. Fifteen percent of apparatus posts carry a quit flag; thirty percent of quit-flagged posts name an apparatus, so seventy percent do not. Neither number says what came first. “Precedes” needs the same writer, an earlier post and a later one, and real posting dates. We have not run that.
Read next. Any introductory treatment of conditional probability and base rates; the error is P(A|B) read as P(B|A).
Dictionary hit rate
A closed list of words and phrases. If a post contains one, it counts as a hit. Fast, auditable, and not a Rosso facet score. We used dictionaries for care talk, mastery talk, optimization, AI/automation mentions, and leaving or strain. On a nursing board the care list mostly fires on ordinary patient talk, so a hit says “this post is about nursing,” not “communion is present.” We keep dictionaries as archive description (how often does AI vocabulary appear at all: under 1% of posts in any year) and draw no year-trend conclusion from them.
Read next. Nelson, Burk, Knudsen, and McCall (2021), “The future of coding,” Sociological Methods & Research, compares hand coding with computer-assisted text analysis. Than et al. (2025) updates that paper for generative models.
Capture year versus posted date
The Internet Archive stamps when it photographed the page, not when the nurse wrote the post. A 2024 snapshot can contain an older thread. We say “capture year” on purpose. It is not an exposure measure and it is not a panel of the same writers.
Read next. The methods page in this bag. No extra citation needed; the error is conceptual, not statistical.
Transfer validation
A codebook built on structured survey stories does not automatically work on messy forum posts. Transfer validation means humans code a fresh sample of those posts (here, 200 already drawn) and you compare the model to those humans before you trust a year contrast. The codebook itself is the one two humans built on the stories. The transfer comparison is the unfinished step: the sheet was drawn and, by PI decision, not coded for this packet. Forum numbers here are that human-built instrument applied by one model to a new text type.
Read next. Törnberg (2024) on domain shift. Ziems et al. (2024) on when computational social science labels fail to travel.
Item response theory (IRT), mentioned so you can ignore it
A homemade 2PL model was fit to unvalidated, binarized draft salience scores, and another to the dictionary hits. Neither is a meaning theta you can take to a client: the salience version has reliability .46, and the dictionary version reproduces the hit count (rank correlation .99) with reliability .20. Both were removed from the board and from this poster page; they are mentioned here only so nobody reconstructs them from the repository and calls them a finding.
Read next. Skip it unless you are the measurement person. Embretson and Reise (2000) is the usual textbook if you are.
A 10-minute path through the citations
| You already know… | Open this |
| Interrater reliability, not ML | Cicchetti and Feinstein (1990); then Törnberg (2024) |
| OLS and p values, suspicious of “n.s.” | Lakens (2017); Wagenmakers (2007) |
| I-O selection / ML hype | Koenig et al. (2023) |
| Meaningful work, not methods | Rosso et al. (2010); Bunderson and Thompson (2009); Bankins and Formosa (2023) |
| Vendor demo tomorrow | Try this quarter and the PSA card above |
Full bibliographic entries sit on the annotated readings page. Visualization sources for how the digital poster was drawn are on the design notes page and are not in the local PDF corpus.