Slides will NOT be shared, so please take notes.
AI/ML for Applied/Empirical Research in Business
Singapore
This Course
Please DO NOT take photos.
We will focus more on methodology rather than applications
We will focus more on intuition rather than math -- many details will be omitted due to the time constraint.
Assumed background: engineering
Prior hands-on experience with AI/ML is assumed.
Building and using models
You know the main AI/ML tasks and model families, including supervised learning, unsupervised learning, and generative AI.
You have worked with data preparation, model training or adaptation, and evaluation on held-out data.
Evaluating a workflow
You recognize overfitting, data leakage, and the role of model selection and validation.
You can assess predictive performance alongside data quality and computational constraints.
We will build on this experience to examine how AI/ML can support applied/empirical research in business.
Assumed background: mathematics
You are comfortable reading the mathematical formulation of a learning problem.
Mathematical and statistical tools
Working knowledge of linear algebra, calculus, probability, and statistics.
Familiarity with random variables, expectations, conditional relationships, regression, and estimation.
The language of learning
Models and parameters; loss functions, empirical risk, and regularization.
Gradient-based optimization, bias–variance trade-offs, and out-of-sample generalization.
Our starting point is applied/empirical research: questions, measurement, identification, and evidence.
What is Applied/Empirical Research?
Applied/Empirical Research
ML for Research
ML as Data Source
ML as Data Source
Every information that is recordable but not numerical can be analyzed with ML to answer business questions
Classified by categories:
Text – NLP
Image – CV
Video – CV
Sound – DL
Biometrics – Bioinformatics
ML as Data Source
ML as Data Source
Machine Learning is used to define new variables from unstructured data
Audio data: Analyzing voice data to gain additional information besides length of calls.
“The Power of Voice: Managerial Affective States and Future Firm Performance,” William Mayew and Mohan Venkatachalam, Journal of Finance, 2012.
“Toward Machines with Emotional Intelligence,” Roslind W. Picard, MIT Media Laboratory report.
ML as Data Source
Machine Learning is used to define new variables from unstructured data
Image data: Analyzing image data to gain additional information besides length of calls.
“How Much Is An Image Worth? An Empirical Analysis of Property’s Image Aesthetic Quality on Demand at AirBNB,” Management Science.
Video data: Analyzing video data to gain additional information of digital ads:
“First Law of Motion: Influencer Video Advertising on TikTok”, Marketing Science
ML as Data Source
If you are using ML to process text data, this reading is good
Dell, Melissa. "Deep learning for economists." Journal of Economic Literature 63.1 (2025): 5-58.
ML as Data Source
Reasons to use ML to understand unstructured data:
Cost reduction and Scalability
Objectivity
Built-in other prediction systems
Problems of using ML to understand unstructured data
Measurement Errors
Interpretation
ML for data processing
Some problem in OLS illustrated
Y = a + b * D + g(X) + epsilon
Outcome
Y can be generated through ML with error
Treatment:
D can be generated through ML with error
Control:
X can be generated through ML with error
X can be selected by ML with error
ML as Prediction Error
Classic Measurement Error and ML Error
The bias can be arbitrary
Better model can mean worse error
Empirical models will not help
Solution
Two general approaches
Standard error re-computation (Instrumental Variable)
De-biasing
But how?
Remember we have training data.
Using training data to learn a better mapping on the error structure
Solution
Solution
ML can also serve as control variables
How to Work With Unstructured Data?
Embedding then Inference
Research Questions
Data Generation Process & Estimation Problem
Step 1: Learn Embedding via Reconstruction
Step 2: Control for the Embeddings
Embedding-then-Inference is Inconsistent
Misaligned Objectives to be Blamed
Debiased Embedding: Neyman Orthogonal Score
Debiased Embedding: Aligning Objectives
Simulation 1: Simple Data Generation Process
Simulation 1: Embedding-then-Inference
Simulation 1: Direct Adjustment with DML
Simulation 1: Comparison
Last, ML as Data Source does not need Causal Inference
ML as Research Methods (Causal Inference)
Causal Inference
Rubin Causal Model
Rubin causal model tells us that we cannot observe the causal effect from observational data (potential outcome is not observable)
We do not observe what happens to Switzerland if people there do not consume chocolate… We also do not observe China/India if people there have a tradition of consuming chocolate…
We have to rely on averages, but averages will give us biases.
We can compare countries with high and low chocolate consumptions, but these countries are inherently different, which creates biases.
When we have random assignments (an experiment), outcome and baseline biases are gone and we can simply compare the treated and control units.
Randomized Control Trials
Average Treatment Effect
Golden Standard
We typically say experiments are golden standards for causal inference.
Why?
The estimator is simple
It achieves the sqrt-n efficiency
It is unbiased in small samples
Basically it suits to tell a causal story
But RCTs can easily be invalid
Go back to our original problem
We have three assumptions essentially to make the above problem super easy:
1. IID of samples
2. SUTVA
3. Random Treatment Assignment
Now, let us try to relax them gradually
What if we do not have randomness?
Aggregating Difference-in-Means Estimators
Aggregating DiM is just Propensity Score
Unconfundedness
Unconfoundedness
ML as Causal Inference
It is hard:
Cross-validation cannot be directly applied to hyper-parameter tuning for causal inference models (Athey and Imbens 2016)
Good performance in predicting the propensity score or outcomes cannot be directly translated into good causal performance (Belloni and Chernozhukov 2014)
Regularization in ML will introduce additional biases (Belloni et al. 2016)
What are the core problems in causal inference?
Non-overlapping/balance, unconfoundedness, power, functional-forms etc.
ML will not magically solve these fundamental problems
ML specifically targeting functional form problems
This is a very fast evolving field (compared to other fields in econometrics)
ML as Causal Inference – Functional-form relaxation
Observational Studies:
Functional form relaxation: ML is used to search for complex functional relationships between features and outcomes. In Causal Inference, we typically assume linear or simple nonlinear relationships. Can we use ML to select better functional forms?
Relax function forms in Regression Analyses:
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., & Newey, W. (2017). Double/debiased/neyman machine learning of treatment effects. American Economic Review, 107(5), 261-65.
Relax function forms in DiD-type of estimators: too many variables we can control, what can we do?
Athey, S., Bayati, M., Doudchenko, N., Imbens, G., & Khosravi, K. (2021). Matrix completion methods for causal panel data models. Journal of the American Statistical Association, 1-15.
ML as Causal Inference – Functional-form relaxation
Observational Studies:
Variable Selection: ML is developed to select the best set of predictors for out-of-sample performance. In causal inference, we often must select variables by hand/experience. ML will help us to solve this problem more systematically
Variable selection in IVs: too many IVs, what can we do?
Belloni, A., Chen, D., Chernozhukov, V., & Hansen, C. (2012). Sparse models and methods for optimal instruments with an application to eminent domain. Econometrica, 80(6), 2369-2429.
Variable selection in Controls: too many variables we can control, what can we do?
Belloni, A., Chernozhukov, V., & Hansen, C. (2014). Inference on treatment effects after selection among high-dimensional controls. The Review of Economic Studies, 81(2), 608-650.
Variable selection in HTEs: too many dimensions, what can we do?
Wager, Stefan, and Susan Athey. "Estimation and inference of heterogeneous treatment effects using random forests." Journal of the American Statistical Association 113.523 (2018): 1228-1242.
Today’s focus
What the paper does?
It provides a general framework for estimating treatment effects using ML methods
The framework requires:
Some regularity condition
ML estimator to converge at n^-1/4 (which you should realize is slower than n^-1/2 and most parametric models, according to delta method, converge at n^-1/2)
The framework outputs:
N^-1/2 – consistent estimator: means that the distribution of the estimator converges in probability to the true estimator with rate n^-1/2.
Many estimators you learnt in parametric world has this properties.
In frequentist world, converge in probability typically also means asymptotically normal, which means you can easily construct standard errors.
Let us look at a very simple example
Partial linear model
Partial linear model
Partial linear model: Naïve Approach
The resulting estimator is biased
Two sources of biases as described by the authors:
Regularization bias:
The ML model cannot perfectly predict the function fast enough according to existing data.
This is fixed by orthogonality moment conditions (this is why the paper is also called Neyman orthogonality)
Overfitting bias:
The ML model may be biased by the existing data compared to the true distribution
This is fixed by sample splitting (very similar to cross validation, which is the key idea in this paper)
Regularization Bias
Regularization Bias
Regularization Bias
Regularization Bias
Regularization Bias
Orthogonality will help us to get rid of b.
The idea is not new. It appears in Robinson (1988) and Frisch-Waugh-Lovell theorem (1930)
More recently, it appears on the authors’ prior paper: Belloni et al. (2014) using lasso
The basic idea is to train two machine learning model and double selecting out the biases. This is called double selection in Belloni et al. (2014).
Frisch-Waugh-Lovell Theorem
Robinson (1988)
FWL and Robinson 1988
DML (2018)
The description in the paper
Now the estimator becomes
Final algorithm for partial linear models
A more general framework
Now, we just cover the illustrative example of paper
Let us look at the real meat in the paper: a general framework in solving this kind of problem
Again, we will use partial linear as an example, and will generalize later
Why does it matter?
We can then change the idea to fit into other problems that is not a simple causal evaluation
A more general framework
Score function
Neyman Orthogonality
DML
DML
The authors show that the estimator through this procedure will
Require:
Some regularity condition
ML estimator to converge at n^-1/4 (which you should realize is slower than n^-1/2 and most parametric models, according to delta method, converge at n^-1/2
Output:
N^-1/2 – consistent estimator: means that the distribution of the estimator converges in probability to the true estimator with rate n^-1/2.
The original data generalization process does not need to be a partial linear model in this more generalized proof!
What is left?
What problems do DML solve?
Deep Learning Based Causal Inference for Large-Scale Combinatorial Experiments: Theory and Empirical Evidence
Outline of the Talk
Introduction: Motivation Example and Potential Solutions
Theory: Debiased Deep Learning and Asymptotics
Empirics: Results with Field Experiment Implementation
A/B Tests are Common Practice for Online Platforms
Multiple A/B Tests on TikTok
Multiple A/B Tests on TikTok
Multiple A/B Tests on TikTok
Solution 1: Linear Addition
Solution 2: Full Factorial Design
Research Questions
Without observing outcomes of all treatment combinations:
Average Treatment Effects
How to estimate and infer the average treatment effect of any treatment combination?
Best Treatment Identification
How to identify the optimal treatment combination with the highest effect?
With minimum changes of the current experimentation platforms
Our Solution and Contributions
Related Literature
Double/de-biased machine learning (DML)
Correct the bias of ML estimators through Neyman orthogonal score functions + cross-fitting
Newey (1994), Chernozhukov et al. (2018, 2022), Farrell et al. (2020, 2021), Athey et al. (2018), Ellickson et al. (2022), Fan et al. (2022), Gordon et al. (2022) etc.
Valid estimation and inference with multiple experiments
Azevedo et al. (2020), Dasgupta et al. (2015), Athey et al. (2021), Pashley and Bind (2019), etc.
(Conjoint Analysis, Green and Rao 1971)
Experiments on online platforms
Evaluating and optimizing the strategies of a large-scale online platform.
Ye et al. (2022), Zeng et al. (2022), Zhang et al. (2020), Cui et al. (2019, 2020), Feldman et al. (2021), Schwartz et al. (2017), Bojinov et al. (2023), Abadie and Zhao (2023) etc.
Outline of the Talk
Introduction: Motivation Example and Potential Solutions
Theory: Debiased Deep Learning and Asymptotics
Empirics: Results with Field Experiment Implementation
Deep Learning Framework: Setup
Goals and Two-stage Method
First stage: Choice of Link Function
How restrictive is this G function?
How restrictive is this G function?
First stage: Deep Learning
First stage: Convergence
Observability and Overlapping Conditions
We need to observe at least (m+2) conditions or (m+1) treatment conditions:
There are m+2 parameters to be estimated
Why do we choose m+2?
In practice, people typically run single experiment + joint holdout experiment, which gives us m+2 conditions including control conditions
Overlapping conditions:
All treatment conditions must be linked to each other through some path.
For example, if m = 3,
We observe (1,0,0), (1,0,1), (0,1,0). Then we cannot run our analysis since the second treatment condition does not link with any other treatment condition.
However, if we observe (0,1,1) instead of (0,1,0), now the second treatment condition is linked to third one through (0,1,1) and the third one is linked to the first one through (1,0,1).
There is a nice graphical representation of the overlapping conditions
Second stage: Naïve Approach
Second stage: Neyman Orthogonality
Intuition behind Neyman Orthogonality
Second stage: Final Estimator
Outline of the Talk
Introduction: Motivation Example and Potential Solutions
Theory: Debiased Deep Learning and Asymptotics
Empirics: Results with Field Experiment Implementation
Field Experiments Setting
We collaborate with one of the top three short-video sharing platforms in the world
Daily Active Users
Monthly Active Users
Ad Revenue per day
Field Experiments Setting
Randomization Check
Ground-Truth ATE
Structured DNN
Benchmarks
Main Result I
DeDL outperforms all benchmarks
A *first empirical validation of DML-type estimator in practice
DeDL Outperforms All Benchmarks
Main Result II
Parametric link function reduces prediction accuracy
De-Bias via Neyman orthogonality
increases ATE estimation accuracy
The benefit of (b) outweights the cost of (a)
Benefit of Neyman Orthogonality
Main Result III
DeDL only *works* when first-stage DNN converges
MAPE Comparison with DNN Training Epoch
Take-away
DeDL framework provides an estimator with provable theoretical guarantees and good empirical performance for analyzing multiple experiments.
DML-type estimator works very well in practice with real experiment data (in contrast to Gordon et al. (2022))
The framework is currently utilized by the platform.
We are trying to open source the framework and you can find preliminary code here:
https://github.com/zikunye2/deep_learning_based_causal_inference_for_combinatorial_experiments
Personalized Policy Learning through Discrete Experimentation: Theory and Empirical Evidence
Platforms Encounter Strategic Decisions Involving Continuous Variables
Imagine you are the manager of the video boosting task team…
Industry Practice: Use Discrete Experiments for Continuous Treatments
Limitation 1: Discrete Treatment May Not Be The Optimal
Limitation 2: Optimal Treatment May Be Heterogeneous
Research Question
Outline of the Talk
Solution Illustration and Contributions
Deep Learning for Policy Targeting Framework
Validation with a Large-Scale Field Experiment
Preview
Theory
Empirics
Outline of the Talk
Solution Illustration and Contributions
Deep Learning for Policy Targeting Framework
Validation with a Large-Scale Field Experiment
Preview
Theory
Empirics
Solution Preview: DLPT Framework
Solution Preview: DLPT Framework
Contributions
Related Literature
Personalized Pricing and Targeting: Leverage on observational/experimental data.
Lu, H., Simester, D., & Zhu, Y. (2025), Yang, J., Eckles, D., Dhillon, P., & Aral, S. (2024), Dubé, J. P., & Misra, S. (2023), Daljord, Ø. et al. (2023), Simester, D., Timoshenko, A., & Zoumpoulis, S. I. (2020),
Experiments on online platforms: Evaluating and optimizing the strategies of a large-scale online platform.
Urban, G. L., Liberali, G., MacDonald, E., Bordley, R., & Hauser, J. R. (2014), Ye et al. (2022), Zeng et al. (2022), Zhang et al. (2020), Feldman et al. (2021), Tucker C. (2014), Tucker, C., & Zhang, J. (2010), etc.
Policy learning with continuous variables: Policy evaluation and optimization.
Kallus & Zhou(2018), Chernozhukov et al. (2019), Athey et al. (2018), etc.
Double/de-biased machine learning (DML): Correct the bias of a plug-in estimator through Neyman-orthogonal score functions.
Newey (1994), Chernozhukov et al. (2018, 2022), Farrell et al. (2020, 2021), Ellickson et al. (2022), Fan et al. (2022), Ye et al. (2023), etc.
Outline of the Talk
Solution Illustration and Contributions
Deep Learning for Policy Targeting Framework
Validation with a Large-Scale Field Experiment
Preview
Theory
Empirics
Set up: DLPT Framework
Key Data Generation Process Assumption
Key Data Generation Process Assumption
Approximation Power of Polynomials
Set Up: Goals
DLPT: Three-Stage Method
Stage 1: Deep Learning
Stage 2: Policy Value Estimation (Neyman Orthogonalized Scores)
Stage 2: Policy Value Estimation (Neyman Orthogonalized Scores)
Stage 3: Policy Learning
What if … Model Misspecification
DLPT: Three-Stage Method
Outline of the Talk
Solution Illustration and Contributions
Deep Learning for Policy Targeting Framework
Validation with a Large-Scale Field Experiment
Preview
Theory
Empirics
Field Experiments Setting
We collaborate with one of the top three short-video sharing platforms in the world
Daily active users
Monthly active users
Total profit per year
Total revenue per year
UGC Platforms Rely on Content Created by Users
Strategies to Encourage Content Creation
My Wallet
Field Experiment
Randomization Check
Empirical Evidence: Cash Reward Incentives Increase Content Creation
Empirical Evidence: Heterogeneous Treatment Effects
Structured Deep Neural Network
Benchmarks
Goal 1: Counterfactual Treatment Evaluation
Goal 1: Counterfactual Treatment Evaluation
DLPT outperforms all benchmarks.
Policy Value Estimation: Cross-Evaluation Strategy
DLPT Outperforms All Benchmarks
Implication 1: Pre-Experiment Design
Goal 2: Personalized Policy Learning
Goal 2: Personalized Policy Learning
(1)DLPT generates the highest profit/lowest regret across policy classes.
(2)Continuous personalized policy improves profit by 6-14%.
Personalized Policy Learning
Policy Classes
Profit Comparison
Implication 2: Selection of Policy Classes
Take-away
High-Dimensional Causal Estimation + Policy Learning
First framework to recover continuous policy values & personalized policies with high dimensional data of discrete A/B test
Strong Theoretical Guarantees with Explicit Bounds on
Approximation error
Estimation error
Policy regret
Empirical Validation
Large-scale field experiment with ~7.4m of users
Shows practical gains from continuity + personalization in real marketing decisions
Practical Impact
The framework is currently utilized by the platform
GitHub link: https://github.com/ZhiqiZhang1229/Policy-Learning-through-Discrete-Experimentation-Theory-and-Empirical-Evidence
A/B Testing is Everywhere
A/B Testing in Research
Faster Decisions of Post-Experiment Analysis
Long Full-Factorial Design => One-time Experiments
Ye, Zikun, Zhiqi Zhang, Dennis Zhang, Heng Zhang, and Renyu Philip Zhang. "Deep-learning-based causal inference for large-scale combinatorial experiments: Theory and empirical evidence." Forthcoming at Management Science.
Finite Discrete Treatments => Continuous Decision Space
Zhiqi Zhang, Zhiyu Zeng, Ruohan Zhan, and Dennis Zhang. "Personalized Policy Learning through Discrete Experimentation: Theory and Empirical Evidence."
Empirical Data Collection => Synthetic Generation
Wang, Mengxin, Dennis J. Zhang, and Heng Zhang. "Large Language Models for Market Research: A Data-augmentation Approach."
Prediction error: step-by-step derivation
ML as Prediction / Optimization
ML as Predictive Decision Making (similar to policy learning)
ML as Optimization
ML to optimize multiple treatments (MAB-type-of paper)
ML to optimize complex MDPs in inventory:
Gijsbrechts, J., Boute, R. N., Van Mieghem, J. A., & Zhang, D. J. (2022). Can Deep Reinforcement Learning Improve Inventory Management? Performance on Lost Sales, Dual-Sourcing, and Multi-Echelon Problems. Manufacturing & Service Operations Management.
Dai, Jim G., and Mark Gluzman. "Queueing network controls via deep reinforcement learning." Stochastic Systems 12.1 (2022): 30-67.
Three stages of DRL research in OM/OR:
Vanilla ML and DRL
ML/DRL for Operations Problems with specific structures
ML/DRL for Operations Problems that select past policies (instead of actions)
The problem of estimation and inference
Sikun Xu, Raphael Thomadsen and Dennis J. Zhang, “Winner’s Curse in Personalized Targeting: Evidence and Solutions”
If you estimate and then optimize, you are probably going to suffer from winner’s curse in evaluation
Personalized Targeting
Winner’s Curse
Winner’s Curse in Targeting One Segment of Homogeneous Consumers
Standard Bootstrapping Can Correct the Winner’s Curse
Standard Bootstrapping Can Correct the Winner’s Curse
Nonsmoothness Slows Down Convergence
Bootstrapping Outperforms Other Proposed Solutions
ML as Structural Model Methods
ML for Structural Estimation
Static structural model (BLP, static games, games on networks, etc.)
Core problems:
Subject to functional assumptions
Slow estimation due to numerical integration
Dynamic structural models (Russ, dynamic games, etc)
Core problems:
Cannot handle large data due to slow estimation and optimization
Subject to very strong functional-form assumptions
Dynamic game models are very restrictive on game structures
Counterfactuals are limited by optimization power
ML as Subjects
ML as Subjects – Human and AI Collaboration
However, many things are still handled by human with algorithms’ help
Human and AI Collaboration
Human and AI Collaboration are important:
In the past, psychology (or behavioral economics) as well as algorithms have been heavily studied, but not the intersection
This topic has huge potential in practice
Human and AI Collaboration in a Process
Research Questions
How to improve human and AI collaboration when humans are recipients of algorithmic decisions?
The Impact of Algorithm-based Work Assignment on Fairness Perceptions and Productivity: Evidence from a Field Experiment, M&SOM.
How to improve human and AI collaboration when humans are users of algorithms?
The Impact of Forced Intervention on AI Adoption, major revision, M&SOM.
How to improve algorithm design with human collaboration data?
Predicting Human Discretion to Adjust Algorithmic Prescriptions: A Large-Scale Field Experiment in Warehouse Operations, Management Science.
ML as Subjects
ML Fairness and Discrimination: https://fairmlbook.org/
Fairness Definitions and Impossibility Results
Test Fairness in AI
Fix Fairness issues in AI
ML and Labor Economics
Eloundou, Tyna, et al. "GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models." arXiv preprint arXiv:2303.10130 (2023).
Acemoglu, Daron, and Pascual Restrepo. "Tasks, automation, and the rise in us wage inequality." Econometrica 90.5 (2022): 1973-2016.
Data Privacy
Transparency in AI
Privacy and Business outcomes:
Goldfarb, Avi, and Catherine E. Tucker. "Privacy regulation and online advertising." Management science 57, no. 1 (2011): 57-71.
Data and ML in IO: Data is now an input into the production and consumption function, which can be priced, traded and incorporated into decisions
Any equilibrium involving algorithms, Data, and other features distinct to AI.
Asker, John, Chaim Fershtman, and Ariel Pakes. "Artificial Intelligence, Algorithm Design and Pricing." In AEA Papers and Proceedings, vol. 112, pp. 452-56. American Economic Association, 2022.
ML as a species: Whether human is only part of the civilization of intelligence?
Butlin, Patrick, et al. "Consciousness in Artificial Intelligence: Insights from the Science of Consciousness." arXiv preprint arXiv:2308.08708 (2023).
Hendrycks, Dan. "Natural selection favors ais over humans." arXiv preprint arXiv:2303.16200 (2023).