Federal Reserve Bank of Chicago
Sep 2026 – Present
Research Assistant, Macroeconomics Team · Chicago, IL
Senior Thesis: Signals in the Transcript
May 2026
Economics, with Honors · Advisor: Dr. Brian Jabarian
- Built a retention prediction pipeline over approximately 12,000 interview transcripts, testing whether their linguistic content carries predictive signal and whether models trained on it predict retention more accurately than the evaluators making the firm's hiring decisions.
- Designed a two-stage residual architecture: transcript models are trained on a structural baseline's residuals rather than the outcome, so their lift isolates the transcript's marginal contribution.
- Estimated two transcript models against the same residual target, each validated on a three-fold expanding window by interview date: (1) a ten-feature linguistic model on XGBoost, with SHAP attribution identifying which features carry retention information; (2) LLaMA-2-7B fine-tuned under QLoRA on a single NVIDIA T4.
- Tested and rejected a judge-leniency instrument after finding that reviewer assignment was not as-good-as-random.
- Primary finding: the firm's revealed decision rule, estimated with propensity scores, places almost no weight on the transcript features; the models recover signal that the existing pipeline leaves unused.
Technologies: Python, HuggingFace Transformers and PEFT, XGBoost, scikit-learn, SHAP, statsmodels, pandas, NumPy, GCP Vertex AI, LaTeX.
[→ Summary: Signals in the Transcript]
University of Chicago Booth School of Business
May 2025 – Jun 2026
Research Assistant (PI: Dr. Brian Jabarian) · Chicago, IL
- Voice AI in Firms (natural field experiment, N ≈ 70,000): built the conceptual framework organizing interview features into four categories (conversation dynamics, interpersonal connection, information quality, language complexity), proposed and justified a 16-variable NLP feature set, and wrote the technical specification for the linguistic style matching index that appears in the paper.
- InterviewAI Learning: implemented SBERT embeddings with UMAP reduction and HDBSCAN clustering to recover roughly 50 coherent interview themes from transcripts, and developed the LLM prompting approach that partitions raw transcripts into structured question and answer pairs with JSON-encoded conversation metadata. Diagnosed memory bottlenecks and migrated the embedding corpus to Cloud Storage for shared access.
- Choice as Signal: drafted the literature review to position the paper's contribution: applicant choice of screening technology as an informative signal. Acknowledged in the paper for research assistance.
- Reviewer misalignment: compared SHAP feature-importance vectors for the human hire decision against those for 90-day retention across 117 evaluators, scoring per-reviewer misalignment and clustering reviewers by pattern to surface signals humans systematically underweight.
- LLM scoring pipeline: encoded an 18-question recruiter rubric into structured prompts, scored candidate samples with Gemini 2.5 Flash, and generated aggregation prompts mapping question-level scores to a hiring decision. Ran the cross-product through the Vertex AI Batch Prediction API, carrying abstentions as null rather than coercing them to a decision, with an online-inference fallback checkpointed per prompt.
Technologies: Python, sentence-transformers, UMAP, HDBSCAN, scikit-learn, SHAP, pandas, NumPy, GCP Vertex AI and Cloud Storage.
[→ Paper: Voice AI in Firms]
[→ Paper: Choice as Signal]
Northwestern University
Jul 2024 – Dec 2024
Research Assistant, Global Poverty Research Lab · Evanston, IL
- Extracted and standardized effect sizes and standard errors from development economics studies on microcredit, business and skills training, financial education, and cash transfers for a World Bank meta-analysis database, the base layer of an evidence library built to make findings comparable across papers.