What we actually did
Study 1 is a two-wave critical-incident panel. Time 1 narratives (March 2016) were dual-coded on 13 Rosso facets. Time 2 (February–March 2017) measured felt meaningfulness, emotional exhaustion, and turnover intentions, the three outcomes the registered tests used, alongside other scales. The interval is about one year (M = 11.58 months). Completers were older and higher in work centrality than noncompleters.
Before any outcome model, we tested whether a language-model panel could replace the human pair. Model instructions were written only on the practice slice, locked (hash-recorded), and scored once on the try-out slice. No set-aside case was scored while any set of instructions could still be chosen. None reached the pass mark. Predictors stayed human. The one later test look is described below.
Studies 2 and 3 use Internet Archive snapshots of allnurses.com, captures January 2016 through 8 September 2026: 7,709 posts, 1,453 threads, 4,443 accounts. The instrument is the Study 1 codebook two humans built on prompted incidents. One model applied it to every post for meaning events (facet × enacted/frustrated/wished × object named × task fate × emotion × attribution) and four post-level flags; 5,435 events, two-thirds of posts with none. A 200-post transfer sheet for a second human coding of that forum prose was drawn and, by PI decision, not coded for this packet. Calls ran at the provider’s default temperature (the lock in the protocol was not applied). Dictionaries cover the full window as archive description. Intervals cluster on thread; account-level resampling was the same or tighter.
Ethics and provenance. Study 1 ran under expedited review by the University of San Francisco IRBPHS (#611); the Time 2 wave was authorized by a modification signed December 27, 2016. The consent form did not address quotation, so no verbatim narrative leaves the private repository; exemplars are paraphrased. Studies 2 and 3 analyzed publicly available forum pages compiled from Internet Archive snapshots: no contact with the site’s servers, no interaction with posters, no private identifiable information, author identifiers hashed at ingest, no post text reproduced. Under 45 CFR 46.102(e) that activity is not human-subjects research and was not submitted for IRB review (PI determination, September 25, 2026). Forum text is handled like respondent microdata: scripts read it, people read aggregates.
What worked
- Locking the drawers. Train / development / test, with development as a single selection pass, kept the substitution test a test: when the prompts missed, there was no path to a quiet retune.
- Dimension-specific bars. Agency and others-directed meaning have human ceilings high enough to test against. Communion and self do not. Writing that down before the run was the difference between a clean verdict and a moving target.
- Humans stayed gold. None of six locked setups reached the agency bar. The outcome models ran on the human codes.
- Reporting the registered estimates as they came in, then screening the rest. The competence–strain association would have been easy to promote. Labeling it exploratory, and putting the ~1,600-test ledger in the supplement, is what makes it usable later.
- Internet Archive rather than a live crawl. The site’s terms forbid systematic retrieval. Snapshots are a completed historical corpus, with capture year stated as the time axis.
- Parser repair as a data problem, not a collection problem. 2024–2025 posts looked empty because later pages store comments in a different HTML tag. Fixing the extractor recovered the years. No new live collection was added.
What we would do differently
- Substitution on this codebook. None of six locked setups reached the agency bar. Agency .750–.798 against .910. Highest .798 is .11 below the bar. The next instrument is a panel judged on outcomes, not on matching one human pair.
- Communal meaning as a confirmatory predictor. Human agreement = .48. We excluded it from the confirmatory family before estimation. Unadjusted Time 2 associations were not robust to unresolved cells.
- Using later URL paths to recover board names. Newer allnurses links drop the forum slug. Hundreds of threads are unlabeled. Student boards leaked into a working-nurse frame.
- Shipping forum rates before the transfer sheet is coded. The codebook is human-built. Comparing it to a second human coding of forum prose is the unfinished gate, not a claim that humans were never in the design.
- Reading a same-post co-occurrence as a lead indicator. Apparatus posts carry quit flags at 15% versus 2%. The reverse conditional is 30%: most quit-flagged posts name no apparatus. One coder, one complaint; not “precedes.”
- Trusting a trend across a site redesign. The site’s page template changed in 2019–2021; posts got shorter and codable events dropped. Any year slope through that band carries the caveat, and the frustration slope does not survive it.
Lessons we would hand to another lab
On measurement
Do not set one acceptance number for “meaning.” Agency and belonging are not the same coding problem. If humans cannot see a dimension, a model fitted to those humans will inherit the blindness and look confident.
On models
A published paper that says models beat crowd workers is not a validation of your codebook. The bar is your human pair, on your facets, on your text type. Forum text is a new validation, not a free extra N.
On outcomes
Concurrent convergence (codes with WAMI in the same wave) is not prospective prediction. We had the first and not the second. Report both or you will overclaim the map.
On time
Capture year is not posted date. Posted date is not AI exposure. Hospital adoption is not this writer. A 2023 cut is the wrong era start; the apparatus talk was there in 2016 at the same rate. Nurses do not say “AI”; they say ratios, metrics, charting. Measure the artefact, not the word. None of this is a causal AI test.
The independent panel (Session 6, Tier 3)
This was specified on 11 June 2026, not invented after the try-out run missed the pass marks. Tier 1 is cells both humans agreed on. Tier 2 is the blinded sample of disagreements. Tier 3 is the panel scored against Time 1 WAMI and Time 2 strain and retention — a criterion no coder, human or model, has seen. Human agreement on communion is .48. A model chosen to match that pair will carry the disagreement forward. A panel that is not selected on human agreement, and is judged on later outcomes, is a different test.
It was not run because ADR 0006 parked all further LLM scoring of the split after every set of instructions missed the human-agreement pass marks and the manuscript had a deadline. The park was a shipping decision. The June design still stands. ADR 0010 puts Tier 3 back on the table. After the manuscript results were fixed, one descriptive look at 65 set-aside cases was taken under a protocol written and hash-recorded beforehand (ADR 0012), on PI decision with coauthor deferral; it scored the two best of the rejected setups and is reported as a 65-case number at try-out precision, not as validation: agency .81 / .80, others-directed .746 / .770, still short of the human bar, salience reliability in the drop band. Same-case human-pair agreement on those 65 was .95 / .84, so the slice was not unusually hard. The pre-set rule for a later single look at the remaining 128 cases (agency and others-directed both ≥ .75) missed on others-directed at .746; the 128 remain sealed and refused by code. Train was the prompt sandbox. Dev is spent. Twenty-nine narratives sit outside the split (eight with Time 2) and can be scored without an unseal; that is a mechanics check, not the powered test.
No respondent text or forum post text is in this packet. The 128-case test remainder stays sealed. If you want the term definitions, use Plain-English glossary.