Short list · why you might open each one

Reading list

Not the manuscript reference list. A short follow-up shelf, with a sentence on each saying what it is good for. Local PDF means the paper sits in this project’s literature folder. “Record only” means the citation is used in the manuscript from a standard bibliographic record; we did not re-open a local PDF for this packet.


Start here if the methods felt foreign

Törnberg, P. (2024). Best practices for text annotation with large language models. Sociologica, 18(2), 67–85. https://doi.org/10.6092/issn.1971-8853/19461

The readable field guide. How to write instructions, how to report errors, why a single accuracy number misleads. Local PDF.

Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30), Article e2305016120. https://doi.org/10.1073/pnas.2305016120

The paper vendors quote. Read it, then read our failure against a locked human bar. Those two facts can both be true. Local PDF.

Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z., & Yang, D. (2024). Can large language models transform computational social science? Computational Linguistics, 50(1), 237–291. https://doi.org/10.1162/coli_a_00502

When labels travel and when they do not. Useful before anyone scores forum text with a survey codebook. Local PDF.

Koenig, N., Tonidandel, S., Thompson, I., Albritton, B., Koohifar, F., Yankov, G., … Newton, C. (2023). Improving measurement and prediction in personnel selection through the application of machine learning. Personnel Psychology, 76(4), 1061–1123. https://doi.org/10.1111/peps.12608

ML in an I-O journal voice: leakage, prediction, and why holdouts exist. Local PDF. Isaac Thompson is a coauthor of that article and of this poster.

Egami, N., Hinck, M., Stewart, B. M., & Wei, H. (2023). Using imperfect surrogates for downstream inference: Design-based supervised learning for social science applications of large language models. Advances in Neural Information Processing Systems, 36.

Why you should not drop machine labels straight into a regression. We specified this and did not run it, because Study 1 used human codes. Local PDF. Read the setup, not the proofs, first.

Agreement and “nonsignificant”

Cicchetti, D. V., & Feinstein, A. R. (1990). High agreement but low kappa: II. Resolving the paradoxes. Journal of Clinical Epidemiology, 43(6), 551–558.

Why we report positive specific agreement instead of leaning on kappa. Record only in this packet; used in the manuscript.

Lakens, D. (2017). Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Social Psychological and Personality Science, 8(4), 355–362. https://doi.org/10.1177/1948550617697177

The cleanest explanation of “not significant” versus “small enough to treat as absent.” Record only; cited in the manuscript TOST supplement.

Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychonomic Bulletin & Review, 14(5), 779–804. https://doi.org/10.3758/BF03194105

Source of the BIC-approximated Bayes factor we report (BF01 = 481.5 on the meaning block). Record only.

Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B, 57(1), 289–300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x

The q = .10 screen on the exploratory search. Record only.

Meaning, helping work, and AI

Rosso, B. D., Dekas, K. H., & Wrzesniewski, A. (2010). On the meaning of work: A theoretical integration and review. Research in Organizational Behavior, 30, 91–127. https://doi.org/10.1016/j.riob.2010.09.001

The 2 × 2 and the open additive, interactive, and curvilinear conjectures (see p. 116 on too many pathways). Local PDF.

Bunderson, J. S., & Thompson, J. A. (2009). The call of the wild: Zookeepers, callings, and the double-edged sword of deeply meaningful work. Administrative Science Quarterly, 54(1), 32–57. https://doi.org/10.2189/asqu.2009.54.1.32

Meaning as infrastructure and as a trap. Record only; confirmed against bibliographic records for the manuscript.

Bankins, S., & Formosa, P. (2023). The ethical implications of artificial intelligence (AI) for meaningful work. Journal of Business Ethics, 185(4), 725–740. https://doi.org/10.1007/s10551-023-05339-7

The ethical map from AI deployment onto meaning. Local PDF.

Flanagan, J. C. (1954). The critical incident technique. Psychological Bulletin, 51(4), 327–358. https://doi.org/10.1037/h0061470

Why we asked for a story, not only a scale. Record only.

Monnot, M. J., & Beehr, T. A. (2014). Subjective well-being at work: Disentangling source effects of stress and support on enthusiasm, contentment, and meaningfulness. Journal of Vocational Behavior, 85(2), 204–218. https://doi.org/10.1016/j.jvb.2014.01.005

Meaningfulness as its own well-being component. That is the Time 2 meaning outcome, distinct from the coded predictor.

Olson, K. D., et al. (2025). Use of ambient AI scribes to reduce administrative burden and professional burnout. JAMA Network Open, 8(10), e2534976. https://doi.org/10.1001/jamanetworkopen.2025.34976

The efficiency premise we are not denying: documentation tools can lower burden. The paradox is what else they remove. Local PDF.

Schellmann, H. (2026). When care becomes code. Scientific American, 334(3), 27.

The UC Davis monitoring episode: a task classed as routine that nurses treated as judgment. Existence proof, not a sample. Verified against the publisher text for the manuscript.

Nelson, L. K., Burk, D., Knudsen, M., & McCall, L. (2021). The future of coding: A comparison of hand-coding and three types of computer-assisted text analysis methods. Sociological Methods & Research, 50(1), 202–237. https://doi.org/10.1177/0049124118769114

Why a dictionary and a human codebook answer different questions. Record only; listed in the project methods bibliography.

Visualization readings are on the design notes page, with notes in design_readings/.