|
on Econometrics |
| By: | Lingwei Kong; Maximilian Osterhaus; Michael Pen |
| Abstract: | This paper develops an inference procedure for average functionals of random-coefficient distributions, such as mean willingness-to-pay and average elasticities, when the distribution is estimated nonparametrically using the penalized fixed-grid estimator of Heiss, Hetzenecker, and Osterhaus (2022). We establish asymptotic normality of the corresponding penalized plug-in estimator centered at the functional evaluated at the penalized pseudo-true value and propose a confidence interval that accounts for the regularization bias. Our method applies to a broad class of linear and nonlinear functionals and allows researchers to use dense grids to reduce approximation bias while maintaining valid inference. Monte Carlo simulations show that the proposed intervals achieve coverage close to the nominal level while remaining informative in finite samples. An empirical application to travel mode demand illustrates that flexible nonparametric specifications can yield economically meaningful differences relative to standard parametric models. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.25416 |
| By: | Fangzhou Yu |
| Abstract: | This paper proposes a robust nonparametric hypothesis test for the existence of heterogeneous treatment effects. We focus on the variance of the Conditional Average Treatment Effect (CATE) as a natural omnibus parameter, where a non-zero variance implies the presence of relevant heterogeneity. Standard inference for this parameter faces a fundamental theoretical challenge. On one hand, evaluating variance components on the same sample leads to null degeneracy, where the asymptotic variance collapses to zero under the null hypothesis of homogeneity, invalidating standard Gaussian inference. On the other hand, decoupling the empirical processes via standard sample-splitting breaks the Neyman orthogonality of the doubly robust scores due to their nonlinear squared loss, which prevents the cancellation of first-order regularization biases. To resolve this challenge, we propose a novel Intra-Fold Sample-Splitting algorithm. By evaluating variance components on mutually disjoint subsamples while coupling them to identical out-of-fold nuisance estimators, our procedure achieves algebraic cancellation of the nuisance biases. We prove this restores consistency and asymptotic normality, and ensures Type I error control. Monte Carlo simulations demonstrate that the proposed test achieves superior size control relative to existing tests while maintaining high power. In an empirical application to the NSW job training program, the test detects significant heterogeneity that traditional nonparametric tests fail to uncover. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.17451 |
| By: | A. Monta\~n\'es; E. Ruiz |
| Abstract: | It is obvious to say that an adequate estimation of the autocorrelation function is central in time series analysis. In this paper, we propose three new robust estimators based on ratios of observations, which offer strong resistance against outliers. While the first estimator, which is based on the median, is not efficient, the second is a Quasi Maximum Likelihood (QML) estimator with better efficiency properties. The third estimator is a plug-in estimator, which does not require numerical optimization and, consequently, is extremely simple from a computationally point of view, having similar efficiency to that of the ML estimator. We derive the asymptotic distribution of the first two estimators, when the true autocorrelations are zero. Furthermore, we also show that the asymptotic distribution of the plug-in estimator is rather close to that of the QML estimator, allowing for inference and, in particular, for the construction of point-wise significance bands for the autocorrelations. Using Monte Carlo simulations, we analyse the finite sample properties of the proposed estimators and compare them with those of the sample autocorrelations and alternative extant robust estimators based on ranks. Although the proposed estimators have larger dispersion than the sample autocorrelations in uncontaminated time series, they are highly robust in the presence of outliers. Also, they have better properties than popular alternative robust estimators based on ranks when estimating autocorrelations of order larger than one. The results are illustrated by estimating the correlogram of daily IBEX35 returns, quarterly US economic growth and monthly US inflation. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.23744 |
| By: | Deborah Kim |
| Abstract: | This article considers the problem of testing sign agreement among a finite number of parameters. This problem arises in empirical settings such as detecting treatment effects with opposite signs across subgroups, outcomes, or time periods, and testing instrument validity for local average treatment effects. For the null hypothesis that the parameters are either all non-negative or all non-positive, I propose two novel tests: a least favorable test and a conditional test. The least favorable test uses a worst-case null critical value, while the conditional test first screens components with large positive or negative estimates and then tests the remaining sign-unresolved components conditional on the screening event. Unlike existing sign agreement tests, both procedures accommodate arbitrary dependence among estimators; in the special case of independent estimators, the critical values depend only on the dimension and testing levels. We show that both tests control asymptotic size uniformly over a large class of nonparametric distributions. Local asymptotic power analysis reveals a tradeoff: the least favorable test is more powerful near boundary configurations where sign restrictions bind, whereas the conditional test is more powerful when some components are well separated from zero. Simulation evidence supports these theoretical predictions in finite samples. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.10294 |
| By: | Viviana Celli; Augusto Cerqua; Guido Pellegrini |
| Abstract: | Spillovers and interference pose fundamental challenges for causal inference, as treatment assigned to one unit may affect the outcome of others, violating the no-interference assumption underlying most empirical strategies. Existing approaches, based on partial interference, exposure mapping, spatial, network, or structural frameworks, typically rely on strong assumptions about interaction structures or require the existence of uncontaminated control units to estimate relevant causal parameters. We revisit this identification challenge within the potential outcomes framework and compare the conditions under which causal effects can be identified using two broad classes of counterfactual methods: control-based counterfactual methods (CBCMs), such as matching and difference-in-differences designs, and forecast-based counterfactual methods (FBCMs), including interrupted time-series and machine learning control methods. We show under which circumstances CBCMs and FBCMs identify average direct and spillover effects. Through simulations and an empirical application, we illustrate the main advantages and limitations of each approach. We show that, in the presence of pervasive or ill-defined spillover effects, CBCMs either cannot be used or entail severe identification concerns, whereas FBCMs can more credibly identify some of the causal parameters of interest, at least in the short term. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.20156 |
| By: | Charisios Grivas; George Kapetanios; Zacharias Psaradakis; Vasilis Sarafidis; Marian Vavra; Alexia Ventouri |
| Abstract: | This paper proposes a nonlinear boosting with multiple testing (BMT) approach to variable selection in high-dimensional generalised linear models with binary responses. At each stage of the BMT procedure, the model is updated by adding only the most significant covariate, conditional on those already selected in previous stages, while taking into account the multiple testing nature of the problem. It is shown that, under the stated conditions, the BMT procedure selects all covariates whose true coefficients are nonzero, and no other covariates, with probability tending to one. Furthermore, the procedure enjoys an oracle property, in the sense that the post-BMT maximum likelihood estimator of the parameters of the model is asymptotically equivalent to an oracle estimator that knows the correct sparse model in advance. Monte Carlo experiments demonstrate that BMT outperforms competing methods, delivering high covariate-selection accuracy and low parameter estimation error. An empirical example illustrates that BMT delivers a predictive model for the probability that U.S. inflation exceeds a given threshold over a 12-month horizon which has very good out-of-sample performance. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.22440 |
| By: | Yuhao Li; Haokun Lu; Xiaojun Song |
| Abstract: | We propose a unified Kernel Minimum Distance (KMD) framework for estimating and testing models defined by conditional moment restrictions. By embedding conditional moments into a Reproducing Kernel Hilbert Space (RKHS), we construct a closed-form $V$-statistic objective function that quantifies the distance from the restrictions. We establish the $\sqrt{n}$-consistency and asymptotic normality of the associated minimum distance estimator. Within this framework, the minimized objective function naturally yields a consistent omnibus specification test. Unlike projection-based methods that require auxiliary nonparametric estimation for Neyman orthogonalization, our test inherently captures the estimation effect via a projected kernel structure. We derive asymptotic properties of the test statistics under the null hypothesis, the alternative hypothesis, and a sequence of local alternatives converging to the null at the parametric rate $n^{-1/2}$. The validity of a computationally simple multiplier bootstrap is established to facilitate inference. Simulation results demonstrate robust finite-sample performance, and the framework is illustrated by analyzing Engel curves using UK Family Expenditure Survey data. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.16605 |
| By: | Juan Estrada; Kim Huynh; David Jacho-Chavez; Leonardo Sanchez-Aragon |
| Abstract: | A novel method to estimate social effect coefficients in the popular so-called linear-in-means regression model in the Social Sciences is presented here that utilizes non-experimental multidimensional network data. The procedure can accommodate social interactions that correlate with the error in the model by making use of a different set of network links among the same observations that are exogenous in the traditional sense. In particular, the full observability of a two-layered multiplex network data structure is assumed here to propose a new Generalized 3-Stage Least Squares (G3SLS) estimator that is consistent, asymptotically normally distributed, and also easy to implement using widely-used existing statistical software because of its closed-form definition. The underlying assumptions are general enough to accommodate common problems with observational data such as measurement error, simultaneity, and unobserved heterogeneity. Monte Carlo exercises confirm the good small sample performance of the proposed G3SLS estimator in these scenarios. An empirical application finds positive and significant peer effects in citations among research articles published in top general-interest journals in economics. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.01421 |
| By: | Binzhi Chen; Annalivia Polselli; Paul S. Clarke |
| Abstract: | Factor structures are central to empirical work in economics and finance, and are usually used to model time-varying unobserved heterogeneity through interactive fixed effects (IFE). Existing IFE estimators rest on low-dimensional and linear specifications in the covariates, assumptions which are increasingly restrictive in applications drawing on rich datasets with controls of unknown functional form. This paper develops a Double Machine Learning estimator for the high-dimensional partially linear panel model with interactive fixed effects (panel DML-IFE). The method combines projection-based defactorisation of the data, in the spirit of Common Correlated Effects (CCE), with a Neyman-orthogonal score function and cross-fitting procedure, and accommodates low-rank factor structures in outcomes and treatments alongside high-dimensional, potentially nonlinear covariate effects estimated by machine learning algorithms. Monte Carlo simulations show that panel DML-IFE outperforms conventional IFE estimator outside the correctly-specified linear case, with bias reduction driven primarily by the time and covariate dimensions. An empirical application to U.S. stock returns shows that several effects documented under linear specifications lose statistical significance once high-dimensional nonlinear confounding and the presence of IFE are jointly accounted for. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.01137 |
| By: | Gregor Steiner; Mark Steel |
| Abstract: | Causal inference is often focused on average effects, which can hide important aspects of the effect distributions. Here we consider the entire posterior effects distribution by estimating full counterfactual outcome distributions. We propose a methodology for inference on counterfactual distributions which builds upon the martingale posterior framework of Fong et al. (2023). This provides a highly flexible approach to estimating densities, distribution functions, and derived quantities such as quantiles, which coherently quantifies the epistemic uncertainty on any target estimand of interest. As the predictive recursions are based on an underlying nonparametric model (a Dirichlet process mixture model), our method naturally inherits robustness with respect to restrictive parametric assumptions. In addition, implementation of our method is typically very fast. This approach can be applied to marginal or conditional counterfactual distributions and is easily extended to an instrumental variables setup. Using the concept of almost conditionally identically distributed random variables, we prove convergence of the martingale posterior inference on the counterfactual outcome distributions for the causal models considered in the paper. We illustrate our approach on both simulated and real data. Using the latter, we investigate the effect of zinc lozenges on common cold duration, the impact of vitamin A supplementation on children's survival rates with one-sided non-compliance (analysed in Imbens and Rubin, 1997a) and the effect of job training (LaLonde, 1986). |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.24143 |
| By: | Qiao, Xinghao; Wang, Zihan; Yao, Qiwei; Zhang, Bo |
| Abstract: | The factor modeling for high-dimensional time series is powerful in discovering latent common components for dimension reduction and information extraction. Most available estimation methods can be divided into two categories: the covariance-based under asymptotically-identifiable assumption and the autocovariance-based with white idiosyncratic noise. This article follows the autocovariance-based framework and develops a novel weight-calibrated method to improve the estimation performance. It adopts a linear projection to tackle high-dimensionality, and employs a reduced-rank autoregression formulation. The asymptotic theory of the proposed method is established, relaxing the assumption on white noise. Additionally, we make the first attempt in the literature by providing a systematic theoretical comparison among the covariance-based, the standard autocovariance-based, and our proposed weight-calibrated autocovariance-based methods in the presence of factors with different strengths. Extensive simulations are conducted to showcase the superior finite-sample performance of our proposed method, as well as to validate the newly established theory. The superiority of our proposal is further illustrated through the analysis of one financial and one macroeconomic datasets. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work. |
| Keywords: | autocovariance;covariance;eigenanalysis;factor strength;reduced rank autoregression;weight matrix |
| JEL: | C1 |
| Date: | 2026–07–27 |
| URL: | https://d.repec.org/n?u=RePEc:ehl:lserod:138585 |
| By: | Cl\'ement de Chaisemartin |
| Abstract: | Difference-in-differences (DID) are sometimes estimated with many pre-treatment periods. In such settings, the observed pre-treatment outcome evolutions provide direct information about the magnitude of shocks that could also occur after treatment. This paper proposes a simple inference procedure that uses those pre-trends as the reference distribution for the post-treatment DID. The procedure is closely related to existing conformal inference procedures, but its DID-specific predictor leads to a distinct identifying restriction. Existing procedures assume parallel trends, while this paper's procedure does not require parallel trends. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.21312 |
| By: | Jos\'e Luis Montiel Olea; Ryan Strong; Amilcar Velez; Zhuoheng Xu; Haomin Yu |
| Abstract: | We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator. Our approximation captures the fact that, in finite samples, the ridge regression estimator trades off bias and variance to reduce estimation and prediction error. Our approximation is based on nonstandard asymptotics where $i)$ we let the estimator's regularization parameter grow proportionally to the sample size; and $ii)$ we treat the population regression coefficients as \emph{local} to the reference vector that defines the estimator's direction of shrinkage. In contrast to other asymptotic approximations in the literature, we allow for general forms of heteroskedasticity and autocorrelation in the data generating process (at the cost of considering a low-dimensional model where the number of covariates is not allowed to grow with the sample size). We use our simple Gaussian approximation to propose two new strategies to select the regularization parameter for the ridge regression estimator. The suggested strategies select the regularization parameter to minimize either average or worst-case excess prediction risk, where risk is computed using our suggested Gaussian approximation. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.02539 |
| By: | Marine Carrasco; Cheikh Nokho |
| Abstract: | This paper proposes methods to estimate and compare asset pricing models in settings with a large number of test assets. Models are specified through a linear stochastic discount factor (SDF). We propose two regularization schemes to extend the Hansen-Jagannathan distance to high-dimensional environments. In addition to stabilizing the inversion of the covariance matrix, the proposed regularizations admit an economic interpretation as relaxing the exact pricing restrictions, thereby accommodating market frictions. We derive the asymptotic properties of the SDF parameter estimator under a double asymptotic framework in which both the cross-sectional and time dimensions grow. These results allow for inference on whether individual factors are priced. We further develop tests for comparing competing asset pricing models under misspeci cation, providing a formal procedure to identify the least misspecified model. The analysis covers both nested and non-nested specifications. An empirical application compares 4 models using a dataset of 647 test portfolios. Cet article propose des méthodes pour estimer et comparer des modèles d’évaluation des actifs dans des contextes où le nombre d’actifs tests est élevé. Les modèles sont caractérisés par leur facteur d’actualisation stochastique (SDF) linéaire. Nous proposons deux méthodes de régularisation permettant d’étendre la distance de Hansen-Jagannathan aux environnements de grande dimension. Au-delà de la stabilisation de l’inversion de la matrice de covariance, les régularisations proposées admettent une interprétation économique comme un relâchement des restrictions exactes d’évaluation des actifs, permettant ainsi de prendre en compte les frictions de marché. Nous établissons les propriétés asymptotiques de l’estimateur des paramètres du SDF dans un cadre de double asymptotique où les dimensions transversale et temporelle tendent toutes deux vers l’infini. Ces résultats permettent de mener des tests afin de déterminer si des facteurs individuels sont valorisés. Nous développons également des tests permettant de comparer différents modèles d’évaluation des actifs en présence d’erreurs de spécification, fournissant ainsi une procédure formelle pour identifier le modèle le moins mal spécifié. L’analyse couvre à la fois les spécifications emboîtées et non emboîtées. Une application empirique consiste à comparer 4 modèles à l’aide de 647 portefeuilles tests. |
| Keywords: | asset pricing models, misspecification, ridge, regularization, tests, Tikhonov, stochastic discount factor, modèles d'évaluation des actifs, spécification erronée, méthode Ridge, régularisation, tests, Tikhonov, facteur d'actualisation stochastique |
| JEL: | C12 G12 |
| Date: | 2026–08–17 |
| URL: | https://d.repec.org/n?u=RePEc:cir:cirwor:2026s-13 |
| By: | Harsh Parikh; Gabriel Levin-Konigsberg; Nilesh Tripuraneni; Dhruv Madeka; Michael I. Jordan; Dean Foster; Dominique Perrault-Joncas; Alexander Volfovsky |
| Abstract: | Randomized controlled trials (RCTs) are fundamental tools for causal inference across technology companies, pharmaceutical research, and federal agencies. While the standard difference-in-means estimator provides unbiased treatment effect estimates, it often lacks precision, particularly when treatment effects are heterogeneous or outcomes exhibit heavy-tailed distributions. Although numerous precision-enhancing methods exist---from covariate adjustment techniques to variance reduction strategies---recent research demonstrates that no single estimator performs optimally across all datasets. Rather than seeking the best estimator for individual RCTs, which risks compromising scientific validity through convenient selection, we propose a principled framework for identifying optimal estimators within families of RCTs based on specific analytical goals. Our approach uses sample splitting to estimate the distribution of evaluation metrics (e.g., mean squared error, regret) across RCT families, enabling systematic comparisons between estimators while maintaining asymptotic guarantees. We demonstrate this framework using a sample of Amazon's Supply Chain Optimization Technology trials and the Strengthening Democracy Challenge dataset (25 interventions). Results reveal that optimal estimators vary significantly by analytical objective: weighted least squares performs best for inference goals, while difference-in-means minimizes regret for decision-making contexts. This work provides actionable guidance for estimator selection while preserving methodological rigor across diverse research applications. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.23254 |
| By: | Xiyu Jiao |
| Abstract: | A common concern in empirical modelling centres around whether estimated regression coefficients are affected by a small set of outlying observations. To conduct outlier robustness checks in practical applications of instrumental variables regressions, the common practice is to run ordinary two stage least squares (2SLS) and remove observations with standardised residuals beyond a chosen cut-off value. Subsequently, the trimmed 2SLS is computed and compared to the original full-sample 2SLS. This paper aims to understand and improve the above heuristic procedure by establishing an asymptotic theory. Specifically, there are three main contributions of the paper. First, the trimmed 2SLS has a positive probability of removing observations even under the null hypothesis where the model contains no outliers. Under this situation, we derive a limiting Normal distribution of the trimmed 2SLS with the asymptotic variance as the ordinary one multiplied by a relative efficiency inflator. Furthermore, a bias correction factor is introduced for the variance estimator of structural errors, which otherwise would be downward biased. Second, a Hausman-type test is constructed to formalize the heuristic procedure of comparing between the two 2SLS estimators. Third, the trimmed 2SLS is a two-step procedure, which can be iterated until a fixed point is reached. The fixed point is shown to have the same first order asymptotics as the Huber-skip M-estimator. Our analysis involves a new class of empirical processes, whose theory would be of independent interest in applied probability. Simulation studies lend support to the asymptotic theory. An empirical illustration to Acemoglu et al. (2019) shows the utility of the proposed method. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.26960 |
| By: | Junjie Li; Yukitoshi Matsushita |
| Abstract: | This paper develops efficient difference-in-differences (DID) estimation under partial interference with a cluster incremental propensity score (CIPS) policy. We define direct and spillover average treatment effects on the treated, establish their identification, and derive their efficient influence functions, from which we construct a cross-fitted estimator. Simulations confirm its finite-sample validity, and an application to China's New Rural Pension Scheme uncovers a significantly negative within-household spillover of pension participation on co-residents' labour income. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.19925 |
| By: | Cash Looi; Ruben Loaiza-Maya; Didier Nibbering |
| Abstract: | Standard multinomial probit (MNP) models specify symmetric latent utility distributions, implying that choice probabilities respond symmetrically to positive and negative covariate shifts of the same magnitude. This restriction is often implausible in empirical choice settings and can lead to misleading elasticity and substitution predictions. We propose a skewed multinomial probit (SMNP) model that captures asymmetric choice responses by specifying a multivariate skew-normal distribution for the latent utilities. The model preserves the flexible substitution patterns of the MNP framework, introduces alternative-specific skewness parameters, and nests the standard MNP model when skewness is zero. Introducing skewness creates identification and computational challenges because the skewness parameters interact with the MNP scale normalization and disrupt the conditional Gaussian updating structure used in Bayesian MNP estimation. We address these challenges through a covariance reparameterization that enforces identification and positive definiteness by construction, interpretable priors on the identified parameter space, and a double data-augmentation scheme that yields a Metropolis-Hastings within Gibbs sampler. Numerical experiments and applications to consumer choice data show that SMNP recovers asymmetric choice responses, improves probabilistic prediction, and produces economically meaningful differences in price elasticities and substitution patterns. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.10336 |
| By: | Qihui Chen; Ka Yan Cheng; Zheng Fang |
| Abstract: | We develop a general framework of identification and estimation for automatic debiased machine learning (DML) where the parameter of interest $\theta_0$ is identified by a moment condition involving a nuisance $\gamma_0$ that may be high dimensional. We establish conditions under which the Riesz representer $\alpha_0$, which is at the core of DML, is identified, and show that the identification occurs precisely when $\alpha_0$ uniquely optimizes a quadratic functional. This characterization enables us to develop a general estimation procedure for $\alpha_0$ that allows for generic $\gamma_0$ including those defined by models with endogeneity and encompasses both classical sieves and modern architectures such as deep neural networks. To improve estimation precision and mitigate the curse of dimensionality, we incorporate shape constraints on $\gamma_0$ by embedding them into a possibly nonlinear parameter space. We illustrate our estimation procedure through simulations and empirical applications. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.24472 |
| By: | Wu, Tianyi; Wang, Tengyao; Samworth, Richard J. |
| Abstract: | In the context of multivariate nonparametric regression with missing covariates, we propose Pattern Embedded Neural Networks (PENNs), which can be applied in conjunction with any existing imputation technique. In addition to a neural net work trained on the imputed data, PENNs pass the vectors of observation indicators through a second neural network to provide a compact representation. The outputs are then combined in a third neural network to produce final predictions. Our main theoretical result exploits an assumption that the observation patterns can be partitioned into cells on which the Bayes regression function behaves similarly, and belongs to a compositional H¨older class. It provides a finite-sample excess risk bound that holds for an arbitrary missingness mechanism, and in combination with a complementary minimax lower bound, demonstrates that our PENN estimator attains in typical cases the minimax rate of convergence as if the cells of the par tition were known in advance, up to a poly-logarithmic factor in the sample size. Numerical experiments on simulated, semi-synthetic and real data confirm that the PENN estimator consistently improves, often dramatically, on standard neural net works without pattern embedding. Code to reproduce our experiments, as well as a tutorial on how to apply our method, is publicly available. |
| Keywords: | deep learning;missing data;nonparametric regression |
| JEL: | C1 |
| Date: | 2026–07–29 |
| URL: | https://d.repec.org/n?u=RePEc:ehl:lserod:139005 |
| By: | Jinglong Zhao |
| Abstract: | We propose a family of control variate estimators for variance reduction in design-based survey sampling and causal inference, with and without interference. In these settings, inverse probability weighting (IPW) estimators are widely used, but may have large variance when sampling, treatment, or exposure probabilities are small. Building on the observation that several common estimators, including the Hajek estimator, the normalized estimator, the augmented inverse probability weighting (AIPW) estimator, and the targeted maximum likelihood estimator (TMLE), all correct the Horvitz-Thompson estimator by canceling part of its randomness, we provide a unified interpretation of these estimators as special cases of a general control variate estimator. We then construct optimal control variates that can reduce the finite sample variance compared to these common estimators. We parameterize the proposed control variates by their bases and characterize the optimal bases through a stochastic optimization formulation. In survey sampling and causal inference without interference, the optimal bases are characterized by leading eigenvectors of matrices that depend on both the design-based sampling structure and the model-based outcome uncertainty. In causal inference under network interference, the optimal bases solve a nonconvex quadratic optimization problem; we provide a $\frac{1}{2}$-approximate solution and an alternating local search heuristic. We apply the control variate estimators to the Swiss Environmental Panel survey data and the Chinese social network data, and conduct extensive simulations to show that the proposed control variate estimators can achieve substantial variance reduction. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.15333 |
| By: | Yuya Shimizu |
| Abstract: | Unstructured data, such as images and text, are increasingly used in empirical economics. Since training machine-learning models on unstructured data is costly, economists often use off-the-shelf pre-trained deep learning models developed by computer scientists to extract embeddings, which are then used as covariates in target economic analyses. Despite the popularity of this practice, its theoretical foundations remain limited. There are two main difficulties. First, the pre-trained model is usually trained on a different dataset and for a different task. Consequently, it is unclear when such a model can be used reliably for the target task. Second, the embedding function is subject to an identification problem, which makes it difficult to analyze the estimation error of the embedding function and its effect on the target task. In this paper, we provide sufficient conditions to overcome these difficulties and derive the convergence rate of machine learning models with pre-trained embeddings. We illustrate the theory through double machine learning applications for estimating parameters of interest, such as partially linear regression with unstructured controls, price elasticity in demand estimation considering the product quality measured by images and text, missing data imputation with unstructured data, and the average treatment effect with unstructured confounders. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.17378 |
| By: | Fangzhou Yu |
| Abstract: | Instrumental-variables estimation increasingly pools many or high-dimensional instruments into a single machine-learned first stage, with rich controls partialled out. The resulting estimand, the partialled-out IV coefficient built from any signal of the instruments, is a signal-weighted average of the heterogeneous effects, which gives an opaque first stage a precise structural meaning. The average is convex whenever a covariance-monotonicity condition holds, and we provide a microfoundation for that condition based on vector monotonicity. With a learned signal, however, the usual debiased moment is not Neyman-orthogonal, and its first-order bias is a drift toward the learner's own signal-weighted average, so naive inference remains valid only for that learner-dependent target. We construct a heterogeneity-robust orthogonal score that restores $\sqrt{N}$ inference on the fixed, learner-invariant target at no efficiency cost, and provide a Hausman-type diagnostic and identification-robust confidence sets. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.17478 |
| By: | Richard Grigorian |
| Abstract: | This paper develops a profiled sieve minimum-distance estimator for a semi-nonparametric differentiated-products demand model with micro-level choice data. Building on Berry and Haile (2024), the estimator uses within-market variation in consumer covariates to recover a flexible consumer-heterogeneity function and market-specific composite intercepts. Excluded price instruments then separate these intercepts into a flexible price-side function and structural demand shocks. The main statistical challenge is that the number of profiled market intercepts grows with the number of markets. I show that, when both the number of markets and the minimum within-market sample size grow, this profiling step is asymptotically negligible and the common structural functions are uniformly consistently estimated. Monte Carlo evidence supports the consistency result and illustrates the value of flexible price-side estimation for counterfactual demand analysis. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.21323 |
| By: | Rametta, Jack T.; Fuller, Sam (Harvard University) |
| Abstract: | Are random forests, the workhorse of supervised machine learning methods in the social sciences, still “good enough” versus new methods that tout dramatic performance benefits? In this article we present a large, diverse Monte Carlo study and existing real-world data benchmarks to compare tree-based methods with TabPFN, a new pretrained transformer foundation model that performs in-context learning over millions of synthetic datasets designed for tabular data. We compare tree-based methods with TabPFN in the popular R-learner framework for conditional average treatment effect estimation. This allows us to assess both predictive model performance and resulting gains for downstream inference. First, our simulations suggest that TabPFN does outperform random forest, achieving near-oracle results for conditional effect estimation. TabPFN’s improvements manifest in the most difficult simulation setups, where the data generating process is complex and there are fewer observations. Second, in real-world data analyses TabPFN performs well, outperforming random forest in some cases, especially as sample dwindles and the number of predictive covariates increases. Our results suggest that tree-based methods are still well suited for social science data, but TabPFN specifically, and the prior data fitted network approach generally, is a strong competitor worthy of consideration. |
| Date: | 2026–07–30 |
| URL: | https://d.repec.org/n?u=RePEc:osf:socarx:g29xc_v1 |
| By: | Nuerxiati Abudurexiti |
| Abstract: | The distribution of a normal mean-variance mixture depends on the law of its positive mixing variable. We compare six parametric mixing laws with a grid nonparametric maximum likelihood estimator under the same determinant identification constraint. The mixing mean $m=\E(Z)$ is estimated and is not fixed at one. A paired block bootstrap is used to compare multivariate holdout log scores. The models that cannot be distinguished from the model with the largest score define a finite ambiguity set. We then consider a cumulative prospect problem on a common portfolio direction. For each model in the set, the NMVM representation gives a scalar projected return and a corresponding prospect-value function of the exposure. The distributionally robust decision maximizes the lower envelope of these functions. We prove existence of a solution, give the candidate points for the piecewise smooth problem, derive a reference-gap scaling result, and construct an interval branch-and-bound certificate for the finite-scenario optimum. In an application to 30 stock returns, the mixture models give higher holdout density scores than the multivariate Gaussian model. Several parametric and semi-parametric models, however, remain in the ambiguity set. The worst-case model is therefore determined at the portfolio optimization stage rather than selected in advance from a point estimate of the holdout score. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.18813 |
| By: | Jaeger, David (University of St. Andrews) |
| Abstract: | Applied economists routinely compare estimates across specifications, observe that they are "similar, " and conclude that their results are "robust.'" This common procedure makes an implicit inferential claim about the range of estimates, but usually does not account for their joint sampling distribution. I formalize informal practice with two bootstrap statistics. The minimum equivalence bound, $R^*_{1-\alpha}$, is the smallest tolerance within which the estimates can be judged equivalent. The range-based equality $p$-value, $p_R$, tests whether the estimates are statistically distinguishable. Together they distinguish failure to detect differences from affirmative evidence of agreement. Simulations show approximately correct size and coverage. Applications to five prominent papers validate some robustness claims while revealing cases in which apparent agreement reflects imprecision rather than stability. A survey of CEPR and NBER affiliates shows that expert judgments aligns with the framework in obvious cases but diverges in intermediate cases. I suggest that $R^*_{.95}$ and $p_R$ be reported whenever multiple specifications are presented as evidence of robustness. |
| Keywords: | robustness, specification sensitivity, equivalence testing, bootstrap inference, joint inference, model uncertainty |
| JEL: | C12 C14 C52 |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:iza:izadps:dp18851 |
| By: | Junchi Shen; Helin Zhao |
| Abstract: | Heavy-tailed diffusion models replace Gaussian noise by a Gaussian variance mixture: denoising Levy probabilistic models (DLPM) take the mixing variables i.i.d. across coordinates, while Student-t EDM shares one mixing variable per sample. Neither has dynamics, yet temporal dependence of the noise amplitude - volatility clustering - is the defining stylized fact of financial returns. We introduce the Denoising Subordinated Probabilistic Model (DSPM), whose mixing vector is a stationary AR(1) chain driven by tempered-stable subordinator increments (the discrete Barndorff-Nielsen-Shephard volatility process) along the data axis. Conditionally on the chain the DDPM machinery survives verbatim; kurtosis and squared-noise autocorrelation are closed-form in the chain parameters, giving an exactly identified, analytically invertible calibration; DDPM, DLPM and Student-t noise are boundary cases of one memory parameter. We then prove a delimiting result: when the denoiser is conditioned on the mixing variables, their law is a nuisance - in the exact-denoiser limit the generated distribution is invariant to it and interventions on the chain do nothing. Experiments confirm both halves: conditioned models match the data's clustering whatever the mixing law, a designed x8 volatility shock moves the envelope by under 13%, while blind models transmit the mechanism exactly as calibrated. Finally, coupling the chain to the data by a variational volatility encoder - trained with the stochastic-volatility likelihood whose log-determinant the simplified denoising loss provably drops - restores control (shock response 3.07 vs. naive 2.83), recovers latent volatility (correlation 0.76), and learns the prior memory toward the true persistence. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.19218 |
| By: | Alexandre Alouadi; Charles-Albert Lehalle |
| Abstract: | Estimating the covariance structure of financial assets typically relies on historical returns, making risk models dependent on noisy and asset-specific time series. We propose the Characteristic-Driven Dynamic Factor Model (CD-DFM), a non-linear latent factor model that instead constructs a representation of the asset cross-section directly from observable firm characteristics, primarily company fundamentals. The learned latent space jointly determines interpretable factor exposures and a forward covariance estimator, and is trained end to end on an objective that combines a Stein covariance loss with a factor reconstruction term, targeting the out-of-sample second moments used in risk management. Because the latent representation, i.e. the encoder depends only on characteristics, previously unseen assets can be embedded at inference time without retraining. Experiments on S&P 500 equities show that CD-DFM produces economically structured latent representations, interpretable factor portfolios, and competitive covariance forecasts despite relying on substantially lower-frequency information than return-based approaches. Among the benchmarked methods, it is the only model that simultaneously combines characteristic-driven representations, factor interpretability, competitive covariance calibration, and zero-shot onboarding of unseen assets. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.24410 |
| By: | Behrooz Moosavi Ramezanzadeh; Arie Beresteanu |
| Abstract: | Partial identification is often set aside in practice because the identification regions it delivers are too wide to be useful, pushing researchers toward strong assumptions that buy point identification at the cost of credibility. We show that a source of information already sitting in most interval-valued datasets can fix this without adding any assumption at all. When an outcome is reported only as an interval---because a data custodian bracketed, top-coded, or formally privatized it to protect respondents---the same custodian typically continues to publish accurate population aggregates of that outcome, precisely because doing so does not compromise any individual record. We develop a framework for exploiting exactly this information: restricting the set of admissible completions of the data to those consistent with a known aggregate, rather than restricting the interval itself, and characterizing the sharp identification region that results for the best linear predictor. The restrictions we study behave in strikingly different ways---some collapse the region by a full dimension, others narrow it while leaving its shape intact. We characterize the geometric effect of each restriction and derive closed-form directional measures of identifying value for the mean and conditional-mean cases. An illustration using interval-valued wages from the Current Population Survey shows that the effect is far from marginal: modest auxiliary information recovers a substantial share of the identifying power usually thought to be lost once an outcome is coarsened. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.21807 |
| By: | Jose, Jibin; Moschini, GianCarlo |
| Abstract: | The production approach recovers markups from output elasticities of flexible inputs and their cost shares in revenue. Recent studies document that markups recovered from different flexible inputs systematically disagree, contradicting the maintained cost-minimization assumption. We show that such conflicting markups arise because standard estimation procedures do not fully exploit the implications of cost minimization. We develop a novel econometric framework that embeds cost-minimization conditions directly into estimation, through cost shares of flexible inputs, while retaining the common Hicks-neutral productivity process. The resulting estimating equation delivers markups that are invariant across flexible inputs by construction. Our findings imply that resolving the conflict in estimated markups across flexible inputs requires only that cost minimization enter the estimation procedure—whether directly, as in our framework, or indirectly via factor-augmenting productivity. These conclusions are illustrated empirically using the Colombian Annual Manufacturing Survey dataset. |
| Keywords: | Research Methods/Statistical Methods |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:ags:aaea26:404719 |
| By: | Dalia Ghanem; Felix Pretis; Daniel Schuurman |
| Abstract: | Understanding the degree to which we are able to adapt to climate change is central to economic assessments of future climate damages. Economists increasingly use comparisons between long differences and fixed effects estimators to measure climate adaptation. We show that such comparisons can be misleading. Neither estimator is consistent for its intended parameter, as both the long-difference (LD) and fixed effects (FE) estimands are weighted averages of the long- and short-run responses to climate and weather. As a result, the difference between the two understates the true extent of adaptation, and the standard test based on this difference --while controlling size -- tends to be substantially underpowered in the settings researchers typically encounter. An empirically-calibrated simulation shows this difference understates adaptation by about 30--80%, depending on the averaging window. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.22028 |
| By: | Aureo de Paula; Elie Tamer |
| Abstract: | Applied econometricians typically model each individual as having fixed outcomes under treatment and control and, in instrumental-variables (IV) settings, fixed treatment decisions under each value of the instrument. This paper asks what changes when outcomes and treatment allocations or choices are stochastic at the individual level. In the model, each individual has a stable (but possibly stochastic) response type consisting of two objects: a treatment choice probability under each state and a potential outcome distribution under each treatment-state pair. These stochastic potential outcomes change the interpretation of some familiar estimators. For instance, in the deterministic IV model, the estimand identifies treatment effect only for compliers-those whose treatment status switches with the instrument. Under stochastic treatment allocation or choice there is no such discrete subgroup: the estimand averages effects over the population, weighing each individual by how much the instrument, policy, or assignment rule moves their probability of treatment. The paper then gives an information-based foundation for stochastic choice, in which individuals act on expected gains given their information. Finally, repeated choices give the stable-kernel formulation empirical content: short panels identify moments of individual treatment probabilities, while long panels identify the joint dependence between treatment effects and the probability movements induced by an instrument or policy. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.21413 |