Tuesday, September 8, 2026

Generating Multivariate Synthetic Data for Asset Prices

Backtesting is a necessary step in strategy development, but it is not sufficient to establish that a strategy is robust. A rigorous validation process is also required.

One inherent limitation of conventional backtesting is that it evaluates a strategy against a single realized historical price path, providing no information about how the strategy might have performed under alternative market trajectories. Techniques such as bootstrapping can help address this limitation by generating alternative price paths. Along this line, Reference [1] proposes a new method for generating realistic synthetic multivariate price paths.

The authors use daily prices and returns for 330 stocks that were members of the S&P 500 at some point between January 2000 and April 2016. To assess whether the synthetic data realistically reproduce the statistical characteristics of financial markets, they examine a broad range of stylized facts and multivariate properties, including fat tails/kurtosis, skewness, autocorrelation, volatility clustering, trend properties, cross-asset correlations, time-varying correlations, and directional similarity among assets.

The authors pointed out,

In this work we have presented an approach to simulate virtual scenarios of multivariate financial data as long as desired. This approach generates artificial asset returns that behave much like the real ones do, as measured by our sanity-check constraints. Virtual scenarios can be simulated, at a low computational cost, involving decades of trading days for hundreds of assets within a given market, and even new artificial assets for that market can be created.

First, in order to define the constraints that artificial asset returns/prices must comply with, the best-known stylized facts have been described and some other properties that account for the cross-asset relationship within a specific market have been introduced. Then, some of the most common approaches found in publicly available toolboxes have been tested in order to check how the simulated asset returns/prices reproduce the observed properties…

On the other hand, our proposed approach allows to generate financial datasets as large as desired while still reproducing volatility clustering and cross-assets relationships (and their changes over time) to a great extent, leading to sets of assets that behave as belonging to the same market…

Basically, their approach is analysis-by-synthesis. First, the authors divide the historical market into upward/downward trends using an equally weighted market index. Within each trend, they estimate time-varying multivariate means and covariance matrices using short rolling windows.

They then generate a stochastic sequence of alternating bull/bear trends and draw returns from the corresponding sequence of time-varying multivariate Gaussian distributions. The generated markets reproduce many characteristics of the real market reasonably closely.

This work makes an important contribution to the growing body of research on synthetic financial data and provides another potentially useful tool for rigorous trading-strategy validation.

Another interesting result of the paper is the use of PCA in reverse. Instead of reducing dimensionality, the authors generate new eigenvectors and project the existing principal components back into asset space, thus creating synthetic stock prices.

Let us know what you think in the comments below or in the discussion forum.

References

[1] Franco-Pedroso, J., Gonzalez-Rodriguez, J., Cubero, J., Planas, M., Cobo, R., & Pablos, F. (2018), Generating virtual scenarios of multivariate financial data for quantitative trading applications. arXiv:1802.01861

Post Source Here: Generating Multivariate Synthetic Data for Asset Prices



source https://harbourfronts.com/generating-multivariate-synthetic-data-asset-prices/

No comments:

Post a Comment