Projects

Latest update: SIOP Leading Edge Consortium 2026 digital goody bag is live — The AI Paradox. Poster, figures, methods, readings, and the quarter kit. QR on the board points here.

A selection of analytics and data-science projects connecting organizational psychology, finance, and AI-driven insight.
For full code and technical details, visit my GitHub profile.

Several of these projects extend my research themes in leadership, well-being, and organizational effectiveness: measuring personality and well-being at scale (2019, MIDUS profiles), keeping selection systems fair (2021), and understanding human skills alongside AI in the future of work (2026).


The AI Paradox (SIOP LEC 2026)

Open the digital goody bag →

Summary figure for The AI Paradox. Study 1, two-wave panel: four Rosso meaning pathways; agency positive specific agreement .93, communion .48; no engineered prompt met the acceptance bar; registered Time 2 null, BF01 = 481.5. Study 2, occupational talk on allnurses.com, captures January 2016 to September 2026, 7,709 posts: nurses named the apparatus, not AI; apparatus talk flat at one post in twenty; control moved, not skill; model codes with no human check.
Study 1 pathways and locked tests beside Study 2 occupational talk, captures Jan 2016 – Sep 2026. Click through to the packet.

Board session packet for Shaping the Future of People Analytics: how optimization can strip the meaning that keeps helping work staffed, and what it takes to measure that meaning without fooling yourself. Two-wave field study, a locked language-model coding pipeline, and ten years of nurses’ forum talk. Human codes stayed on the map; no engineered prompt met the acceptance bar; no prompt was chosen on test. In the forum, nurses named the apparatus (ratios, metrics, charting), not AI, and what moved was control over the work, not skill.

Monnot, M. J., & Thompson, I. (2026, September 30–October 1). The AI paradox: How optimization may undermine meaningful work [Poster presentation]. Society for Industrial and Organizational Psychology Leading Edge Consortium, Baltimore, MD, United States. https://mjmonnot.github.io/lec-2026/

🔗 Live packet · Digital poster · Print deck (PPTX)


SIOP Machine Learning Competitions

View on GitHub →

A year-over-year, reproducible collection of solutions and teaching cases for the SIOP Machine Learning Competitions.

2026 — Automated Meta-Analytic Coding (team: One Hot Key)

Six-gate cascading pipeline from the SIOP 2026 presentation: PDF acquisition, layout extraction, regex plus phi4 classifier, vision fallback, structured LLM extraction, imputation.
Six gates from PDF to correlation table. Each gate only fires when the cheaper one before it fails.

An end-to-end pipeline that extracts zero-order Pearson r correlations directly from published I-O psychology PDFs, using a four-tier cascade (pdfplumber → Docling table ML → qwen2.5-VL vision model → regex + phi4) running entirely on local models. Built as a solo-plus-AI-agents experiment — a one-person team competing against teams of researchers and graduate students. Dev-set MSE 0.013641 (6th of 24); test set submitted April 2026.

🔗 Project folder → · SIOP 2026 deck (PDF) · SIOP 2026 presentation video (MP4)
Relevant resources: Docling (document & table extraction) · PRISMA — systematic review & meta-analysis reporting

2019 — Personality Prediction from Text (Post-Hoc Winning Solution)

Four-step role-play questionnaire pipeline: read the five text answers, role-play the persona, answer 30 BFI-2 items in character, reverse-score and aggregate to OCEAN.
Role-play scoring: the model answers a BFI-2 questionnaire as the respondent, then the items are scored like any inventory.

A post-hoc solution to the 2019 competition — predicting Big Five trait scores from five short open-ended responses — that beats the original leaderboard by a wide margin under a strict, leakage-safe protocol (fit on Train only, select on Dev, touch the private Test once).

  • Private-Test mean Pearson r 0.3215 vs. the 2019 first-place 0.26021 — +0.061 (~23% relative), roughly 2× the entire original top-four spread
  • Stacked generalization: zero-shot LLM extractors (multi-prompt trait scoring, a second-judge model, behavioral subfeatures, and a role-play BFI-2 questionnaire) combined with embedding-SVR, TF-IDF, and psycholinguistic bases under a per-trait Ridge meta-learner
  • Result sits at or above the 2025–2026 published frontier for personality inference from short text (e.g., Piastra & Catellani, 2025; Zhu et al., 2025); full write-up, negative results, and literature review included
  • Directly extends my measurement research: construct validity, honest evaluation, and personality assessment at scale

🔗 Project folder → · Poster (PDF) · Presentation deck (PDF) · Presentation video (MP4)
Relevant resources: Solution write-up (SOLUTION.md) · International Personality Item Pool (IPIP)


Archived / Post-Hoc Competition Reconstructions

2024 — Evaluating LLMs Across Four I-O Tasks (Post-Hoc Reconstruction)

  • One unified prompt-engineering harness spanning all four 2024 tasks — empathy, interview generation, item clarity, and fairness
  • Shared format → call → parse flow with task-specific adapters, structured (constrained JSON) outputs, and similarity-based few-shot selection
  • Final scorecard pitting a single unified pipeline against four separately hand-tuned winning teams, task by task
  • Full end-to-end sweep on synthetic inputs recorded test composite 0.817 (dev 0.814); not comparable to winner scores on official competition data

🔗 Project folder → · SIOP 2024 retrospective deck (PDF) · SIOP 2024 retrospective video (MP4)
Relevant resources: Original 2024 SIOP ML Competition · Sentence-Transformers (SBERT)

2023 — Decision Making from Text

  • Predicting assessment-center “decision making” ratings from open-ended text
  • End-to-end pipeline: validate → preprocess → features → train → evaluate → fairness audit
  • Transparent TF-IDF + Ridge baseline with a transformer-ready (SBERT) template

🔗 Project folder →
Relevant resources: Guidelines for Assessment Center Operations · Sentence-Transformers (SBERT)

2021 — Fairness-Aware Selection Pipeline (Teaching Case)

  • Decomposition of accuracy vs. adverse impact tradeoffs
  • Alignment with professional standards for employee selection
  • Designed as a teaching and practitioner case

🔗 Project folder →
Relevant resources: Uniform Guidelines on Employee Selection Procedures · SIOP Principles for Personnel Selection

(Subsequent years will be added as independent modules.)


Methods & Tooling

Python · Pandas · NumPy · scikit-learn pipelines · cross-validation · regularization and ensembles · stacked generalization (out-of-fold meta-learning) · zero-shot LLM feature extraction (Anthropic Claude) · sentence embeddings (E5, SBERT) · PDF extraction (PyMuPDF · pdfplumber · Docling) · local language and vision models (phi4 · qwen2.5-VL, served via Ollama) · model diagnostics · fairness metrics · reproducible GitHub workflows


Afloat or Adrift: Latent Personality Profiles & Future-of-Work Skills (MIDUS)

View on GitHub →

Figure 1A from the MIDUS preprint: latent state means on anchored factor scores for four personality profiles across neuroticism, extraversion, openness, agreeableness, conscientiousness, and agency. Resilient 36 percent, Distressed 30 percent, Reserved 29 percent, Antagonistic 5 percent.
Figure 1A from the preprint: the four replicated profiles on six anchored trait scores.

Overview:
A fully reproducible, longitudinal study of person-centered Big Five personality profiles and how they relate to future-of-work skills across midlife, using the MIDUS (Midlife in the United States) national panel (N = 7,108 over ~20 years) with independent replication in the MIDUS Refresher (N = 3,577). Where the SIOP 2019 project predicts traits from text, this project asks what trait configurations mean — and, crucially, whether people move between them over two decades — connecting directly to my research on well-being and meaningful work. Framed around self-determination theory and the psychological resources workers need to develop and retain AI-era skills.

Built With:
R · latent profile analysis (LPA) and latent transition analysis via a joint latent Markov model · BCH / 3-step outcome modeling · Mplus confirmation · GitHub Actions CI · devcontainer for a reproducible environment

Key Findings:

  • Four replicated profiles: Resilient (36.4%), Distressed (29.5%), Reserved (29.0%), and Antagonistic (5.1%). Resilient members reported the highest psychosocial skills; Antagonistic members reported the highest income, prestige, and analytic performance.
  • An honest null on incremental validity: profile membership added no predictive power beyond continuous traits.
  • The Distressed profile shrank from 29.5% to 7.2% across two decades through two distinct pathways — recovery (movement to Resilient) and disengagement (movement to Reserved).
  • Leaving the Distressed profile was associated with lower health-related lost productive time (recovered movers −$3,049 per worker-year, 95% CI [−$5,038, −$1,060]).
  • Purpose in life predicted recovery- versus disengagement-oriented transitions (OR = 1.23 per SD, 95% CI [1.01, 1.50]).
  • Event-sampled diary data (N = 2,314) corroborated the profile interpretations.

🔗 Read the preprint (PsyArXiv) →
Manuscript under review.


AI Bubble Pressure Score (AIBPS)

View on GitHub → · Live dashboard →

Summary figure for the AI Bubble Pressure Score: six pillars (market, credit, capex, infrastructure, adoption, sentiment) feed a four-step pipeline of ingest, normalize, aggregate, and interpret, producing a 0 to 100 composite with low, healthy, frothy, and critical bands.
Six pillars, one 0–100 composite. Click through to the repository.

Overview:
An ongoing project analyzing sentiment, valuation, and market momentum to estimate “bubble pressure” in the AI sector. The AIBPS integrates multiple data layers—equity performance, ETF flows, and public sentiment—to quantify how narrative intensity and capital inflows co-evolve across AI-related assets.

Built With:
Python · Pandas · NumPy · Matplotlib · scikit-learn · GitHub Actions · CSV/REST data pipelines

Key Features:

  • Automated ingestion of market and sentiment data
  • Rolling normalization (z-scores, percentiles)
  • Composite pressure index tracked over time
  • Automated updates via GitHub Actions

Use Cases:

  • Identify when enthusiasm and capital concentration approach “hype cycle” territory
  • Compare AI-related funds and semiconductor equities
  • Demonstrate applied analytics for investment or strategic contexts

AI Hyperscaler Market-Cap Race

Final frame of the AI hyperscalers market-cap bar chart race, September 2026: NVIDIA leads above 5 trillion dollars, followed by Alphabet 4,162 billion, Microsoft 3,833, Amazon 2,686, TSMC 2,337, Meta 1,915, and AMD 1,028.
Final frame, September 2026. Click through to run the race.

A companion data-visualization piece: an animated D3.js bar chart race tracking the monthly market capitalization of leading AI hyperscalers and infrastructure firms over time, auto-updated through a GitHub Actions pipeline with no backend.

🔗 Live visualization → · View on GitHub →


Future Additions

Additional projects will be added over time, including:

  • Leadership and coaching analytics
  • Employee well-being dashboards building on the MIDUS profile work
  • Applied ML and measurement pipelines
  • Visualization tools for organizational and market psychology

Explore All Repositories

🔗 Browse all public GitHub repositories →