Showing posts with label hft. Show all posts
Showing posts with label hft. Show all posts

Sunday, June 9, 2013

Nanex: visualizing zillions of trades in a half-second

By way of Naked Capitalism, I learned of the video below produced by Nanex, the market data analysis company in Chicago. It runs for just under 6 minutes and shows all the quotes for Johnson & Johnson stock racing among a network of exchanges over just one half second of real time. Watch the whole thing and then remember -- this is just one half second (the clock, bottom middle, goes up in increments of milliseconds). Below the video, some explanation of what you're seeing from Nanex.




We got the idea after realizing, in face to face meetings with them [the CFTC], they didn’t understand market structure or the importance of latency and the consolidated feed. That was several years ago. We still aren’t sure if they get it, or are just playing dumb.

The bottom box (SIP) shows the National Best Bid and Offer. Watch how much it changes in the blink of an eye.

Watch High-Frequency Traders (HFT) at the millisecond level jam thousands of quotes in the stock of Johnson and Johnson (JNJ) through our financial networks on May 2, 2013. Video shows 1/2 second of time. If any of the connections are not running perfectly, High Frequency Traders can profit from the price discrepancies that result. There is no economic justification for this abusive behavior.

Each box represents one exchange. The SIP (CQS in this case) is the box at 6 o’clock. It shows the National Best Bid/Offer. Watch how much it changes in a fraction of a second. The shapes represent quote changes which are the result of a change to the top of the book at each exchange. The time at the bottom of the screen is Eastern Time HH:MM:SS:mmm (mmm = millisecond). We slow time down so you can see what goes on at the millisecond level. A millisecond (ms) is 1/1000th of a second.

Note how every exchange must process every quote from the others — for proper trade through price protection. This complex web of technology must run flawlessly every millisecond of the trading day, or arbitrage (HFT profit) opportunities will appear. It is easy for HFTs to cause delays in one or more of the connections between each exchange.

Thursday, May 2, 2013

The state of high frequency trading -- interview with Dave Cliff

This is a very short interview with Dave Cliff of the University of Bristol who knows more about high frequency trading, the good and the bad, than just about anyone else. Short take: the wild west period may be over, the potential for huge profits has been arbitraged away, HFT now looks more like part of the fixed and settled landscape than something really new. 

Friday, November 30, 2012

The infallible portfolio?

In this short essay, Ricardo Fenholz of Columbia University makes what seems (to me at least) to be a rather incredible claim: that it's relatively easy to construct a portfolio that is guaranteed to outperform the S&P 500 over one year (or any other interval you like) and also has a limited downside during that year. The idea is the take the S&P 500 Index and tweak it a little, creating a portfolio with less weight in stocks with higher capitalization, and more weight in those with less capitalization, and presto -- you have something guaranteed to outperform the S&P 500, he claims. Is that possible? That easy? Here's a little more detail:
To understand how this works, let’s consider the S&P 500 U.S. stock index. Suppose that we wish to invest some money in S&P 500 stocks for one year. Currently, Apple has a total market capitalization of roughly $500 billion, making it the largest stock in the S&P 500 and equal to approximately 4% of the total capitalization of the entire index. Suppose that we believe it is very unlikely or impossible that either Apple or any other corporation’s capitalization will be equal to more than 99% of the total S&P 500 capitalization for this entire year during which we plan to invest. As long as this turns out to be true, then it is actually pretty simple to construct a portfolio containing S&P 500 stocks that is guaranteed to outperform the S&P 500 index over the course of the year and that has a limited downside relative to this index. In essence, we can construct a portfolio that will never fall below the value of the S&P 500 index by more than, say, 5% and that is guaranteed to achieve a higher value than the S&P 500 index by the end of the year.[1]
This is not a trivial proposition. If we combine a long position in this outperforming portfolio together with a short position in the S&P 500 index, then we have a trading strategy that requires no initial investment, has a limited downside, and is guaranteed to produce positive wealth by the end of the year. According to standard financial theory, this should not be possible.[2] Furthermore, the assumptions that guarantee that our portfolio will outperform the S&P 500 index appear entirely reasonable. After all, not for one day in the more than 50-year history of the S&P 500 has one corporation’s market capitalization come anywhere close to equaling even 50% of the total capitalization of the market. A 99% share of total market capitalization would essentially amount to there being only one corporation in the entire U. S. for an entire year. This seems like neither a likely outcome nor one that investors should take seriously when constructing their portfolios.
What does a portfolio made up of S&P 500 stocks that is guaranteed to outperform the S&P 500 index look like? There are many different ways in which such a portfolio can be constructed, but one feature common to all such portfolios is that relative to the S&P 500 index itself, they place more weight on those stocks with small total market capitalizations and less weight on those stocks with large total market capitalizations. The weight that an index such as the S&P 500 places on each individual stock is equal to the ratio of that stock’s total market capitalization relative to all stocks’ total market capitalizations taken together. In the case of Apple, then, the S&P 500 index would place a weight of roughly 4% in this individual stock while those portfolios that use HFT to outperform this index would instead place a weight of less than 4% in Apple stock.
The only condition for this to work, he suggests, is that the assumption that no stock in the market comes to dominate the market in the sense that its market capitalization comes to be a high fraction of that of the entire market. This is, as he notes, a fairly weak assumption, although the weaker you make the assumption, the longer the time interval over which this idea apparently works.

Now, I'm not doubting the veracity of this claim. I'm just stunned that such a simple recipe could work, and can't see the intuition behind it. What if the high cap stocks happen to perform brilliantly next year, relative to the lower cap stocks? Wouldn't this portfolio with underweighted high cap stocks then underperform the S&P Index? I've had a quick look at the paper Fernholz references as a detailed support of his claim, and there he explains the conditions for the theorem to hold in slightly different terms:
The conditions mandate, roughly, that the largest stock have "strongly negative" rate of growth, resulting in a sufficiently strong repelling drift away from an appropriate boundary; and that all other stocks have "sufficiently high" rates of growth.
That sounds very different from the quite plausible assumption about no market dominance of a single stock. Indeed, this seems like saying that, if one assumes that large cap stocks will perform poorly, and small caps stock better, then we can build a portfolio guaranteed to outperform the S&P 500 Index by weighting small cap stocks more heavily. Isn't that like assuming we know the future?

But maybe I'm wrong. I'd be interested in the thoughts of others. The paper is quite dense and light on intuitive discussion of the logic. Fernholz suggests that perhaps the existence of these superior portfolios -- which require continuous rebalancing by high-frequency buying and selling of many stocks -- explains some of the very high profits consistently earned by quantitative high-frequency hedge funds such as Renaissance Technology's Medallion Fund. I think I find more convincing the analysis of Lo and Khandani which seemed to suggest that much of the performance of quant hedge funds over the past decade or so can be accounted for by fairly vanilla long short equity strategies, with increasing use of leverage in the mid 2000s (used to maintain high reported earnings even as raw earnings fell off due to competition).

Monday, February 13, 2012

Approaching the singularity -- in global finance

In a new paper on trends in high-frequency trading, Neil Johnson  and colleagues note that:
... a new dedicated transatlantic cable is being built just to shave 5 milliseconds off transatlantic communication times between US and UK traders, while a new purpose-built chip iX-eCute is being launched which prepares trades in 740 nanoseconds ...
This just illustrates the technological arms race underway as firms try to out-compete each other to gain an edge through speed. None of the players in this market worries too much about what this arms race might mean for the longer term systemic stability of market; it's just race ahead and hope for the best. I've written before (here and here) about some analyses (notably from Andrew Haldane of the Bank of England) suggesting that this race is generally increasing market volatility and will likely lead to disaster in one form or another. 

We may be getting close. If Johnson and his colleagues are correct, the markets are already showing signs of having already made a transition into a machine-dominated phase in which humans have little control.

Most readers here will probably know about futurist Ray Kurzweil's prediction of the approaching "singularity" -- the idea that as our technology becomes increasingly intelligent it will at some point create self-sustaining positive feedback loops that drive explosively faster science and development leading to a kind of super-intelligence residing in machines. Humans will be out of the loop and left behind. Given that so much of the future vision of computing now centers on bio-inspired computing -- computers that operate more along the lines of living organisms, being able do things like self-repair and adaptation, true reproduction, etc, -- it's easy to imagine this super-intelligence again ultimately being strongly biological in form, while also exploiting technologies that earlier evolution was unable to harness (superconductivity, quantum computing, etc.). In that case -- again, if you believe this conjecture has some merit -- it could turn out ironically that all of our computing technology will act as a kind of mid-wife aiding a transition from Homo sapiens to some future non-human but super-intelligent species.

But forget that. Think singularity, but in the smaller world of the markets. Johnson and his colleagues ask the question of whether today's high-frequency markets are moving toward a boundary of speed where human intervention and control is effectively impossible:
The downside of society’s continuing drive toward larger, faster, and more interconnected socio-technical systems such as global financial markets, is that future catastrophes may be less easy to forsee and manage -- as witnessed by the recent emergence of financial flash-crashes. In traditional human-machine systems, real-time human intervention may be possible if the undesired changes occur within typical human reaction times. However,... in many areas of human activity, the quickest that someone can notice such a cue and physically react, is approximately 1000 milliseconds (1 second)
Obviously, most trading now happens much faster than this. Is this worrying? With the authors, let's look at the data.

In the period from 2006-2011, they found (looking at many stocks on multiple exchanges) that there were about 18,500 specific episodes in which markets, in less than 1.5 seconds, either 1. ticked down at least 10 times in a row, dropping by more than 0.8% or 2. ticked up at least 10 times in a row, rising by more than 0.8%. The figure below shows two typical events, a crash and a spike (upward), both lasting only 25 ms.



Apparently, these very brief and momentary downward crashes or upward spikes -- the authors refer to them as "fractures" or "Black Swan events" -- are about equally likely. And they become more likely as one goes to shorter time intervals:
... our data set shows a far greater tendency for these financial fractures to occur, within a given duration time-window, as we move to smaller timescales, e.g. 100-200ms has approximately ten times more than 900-1000ms.
But they also find something much more significant. They studied the distribution of these events by size, and considered if this distribution changes when looking at events taking place on different timescales. The data suggests that it does. For times above about 0.8 seconds or so, the distribution closely fits a power law, in agreement with countless other studies of market returns on times of one second or longer. For times shorter than about 0.8 seconds, the distribution begins to depart from the power law form. (It's NOT that it becomes more Gaussian, but it does become something else that is not a power law.) The conclusion is that something significant happens in the market when we reach times going below 1 second -- roughly the timescale of human action.

Ok. Now for the punchline -- an effort to understand how this transition might happen. In my last blog post I wrote about the Minority Game -- a simple model of a market in which adaptive agents attempt to profit by using a variety of different strategies. It reproduces the realistic statistics of real markets, despite its simplicity. I expect that some people may wonder if this model can really be useful in exploring real markets. If so, this new work by Johnson and colleagues offers a powerful example of how valuable the minority game can be in action.

Their hypothesis is that the observed transition in market dynamics below one second reflects "a new fundamental transition from a mixed phase of humans and machines, in which humans have time to assess information and act, to an ultrafast all-machine phase in which machines dictate price changes."  They explore this in a model that...
...considers an ecology of N heterogenous agents (machines and/or humans) who repeatedly compete to win in a competition for limited resources. Each agent possesses s > 1 strategies. An agent only participates if it has a strategy that has performed sufficiently well in the recent past. It uses its best strategy at a given timestep. The agents sit watching a common source of information, e.g. recent price movements encoded as a bit-string of length M, and act on potentially profitable patterns they observe.
This is just the minority game as I described it a few days ago. One of the truly significant lessons emerging from its study is that we should expect markets to have two fundamentally distinct phases of dynamics depending on the parameter α=P/N, where P is the number of different past histories the agents can perceive, and N is the number of agents in the game. [P=2M if the agents use bit strings of length M in forming their strategies]. If α is small, then there are lots of players relative to the number of different market histories they can perceive. If α is big, then there are many different possible histories relative to only a few people. These two extremes lead to very different market behaviour.

Johnson and colleagues suggests that the transition between these regimes is just what shows up in the statistics around the one second threshold. They first argue that the regime for large α (many strategies per agent) should be associated with the trading regime above one second, where both people and machines take part. Why? As they suggest,
We associate this regime (see Fig. 3) with a market in which both humans and machines are dictating prices, and hence timescales above the transition (>1s), for these reasons: The presence of humans actively trading -- and hence their ‘free will’ together with the myriad ways in which they can manually override algorithms -- means that the effective number (i.e. α > 1). Moreover α > 1 implies m is large, hence there are more pieces of information available which suggests longer timescales...  in this α > 1 regime, the average number of agents per strategy is less than 1, hence any crowding effects due to agents coincidentally using the same strategy will be small. This lack of crowding leads our model to predict that any large price movements arising for α > 1 will be rare and take place over a longer duration – exactly as observed in our data for timescales above 1000ms. Indeed, our model’s price output (e.g. Fig. 3, right-hand panel) reproduces the stylized facts associated with financial markets over longer timescales, including a power-law distribution.
What they're getting at here is that crowding in the space of strategies, by creating strong correlations in the strategies of different agents, should tend to make large market movements more likely. After all, if lots of agents come to use the very same strategy, they will all trade the same way at the same time. In this regime above one second, with humans and machine, they suggests there shouldn't be much crowding; the dynamics here do give a power law distribution of movements, but it is what is found in all markets in this regime.

In contrast, they suggest that the sub one second regime should be associated the the α < 1 phase of the minority game:
Our association of the α < 1 regime with an all-machine phase is consistent with the fact that trading algorithms in the sub-second regime need to be executable extremely quickly and hence be relatively simple, without calling on much memory concerning past information: α < 1 regime with an all-machine phase is consistent with the fact that trading algorithms in the sub-second regime need to be executable extremely quickly and hence be relatively simple, without calling on much memory concerning past information: Hence M will be small, so the total number of strategies will be small and therefore... α < 1. Our model also predicts that the size distribution for the black swans in this ultrafast regime (α < 1) should not have a power law since changes of all sizes do not appear – this is again consistent with the results in Fig. 2.
And...
Our model undergoes a transition around α = 1 to a regime characterized by significant strategy crowding and hence large fluctuations. The price output for α < 1 (Fig. 3, left-hand panel) shows frequent abrupt changes due to agents moving as unintentional groups into particular strategies. Our model therefore predicts a rapidly increasing number of ultrafast black swan events as we move to smaller α and hence smaller subsecond timescales – as observed in our data.
 The authors go on to quantify this transition in a little more detail. In particular, they calculate in the simple minority game model the standard deviation of the price fluctuations. In the regime α < 1 this turns out to be roughly proportional to the number N of agents in the market. In contrast, it goes in proportion only to the square root of N in the α > 1 regime. Hence, the model predicts a sharp increase in the size of market fluctuations when entering the machine dominated phase below one second.

The paper as a whole takes a bit of time to get your head around, but it is, I think, a beautiful example of how a simple model that explores some of the rich dynamics of how strategies interact in a market can give rise to some deep insights. The analysis suggests, first, that the high frequency markets have moved past "the singularity," their dynamics having become fundamentally different -- uncoupled from the control, or at least strong influence, of human trading. It also suggests, second, that the change in dynamics derives directly from the crowding of strategies that operate on very short timescales, this crowding caused by the need for relative simplicity in these strategies.

This kind of analysis really should have some bearing on the consideration of potential new regulations on HFT. But that's another big topic. Quite aside from practical matters, the paper shows how valuable perspectives and toy models like the minority game might be.

Tuesday, January 24, 2012

Markets -- increasingly complex dynamics over the past decade

Didier Sornette is among the most creative scientists I know, and always seems to come up with an approach to problems that is more or less orthogonal to what anyone has done before. In a paper just out (as a preprint), he and Vladimir Filimonov offer a really novel analysis on the old question about whether market movements are caused by A. external influences such as news (exogenous causes) or B. influences internal to the market itself such as emotions, avalanches of belief and opinion, etc. (endogenous causes). This matter, of course, touches directly on the infamous efficient markets hypothesis, which insists on interpretation A (all A, no B).

I've written before (here and here, for example) about various studies trying to match up news feeds with big market moves to see if the latter can be explained by the former. Generally, the evidence suggests no, implying some mixture of A and B. Sornette and Filimonov now take a very different approach, which is an attempt to use mathematics to directly measure how much of the dynamics of a time series can be attributed to endogenous, internal causes. The mathematical technique is itself interesting. If it can be trusted, then the results suggest that markets in the past decade have become much more strongly driven by internal, endogenous dynamics than they were before. As the authors point out, this could well reflect the explosion of algorithmic trading, as computers interact with one another in lots of complex feedback loops.

The authors envision their technique as a device for measuring the amount of "reflexivity" in the market, referring to the term used by George Soros to describe how human perceptions and misperceptions interact in the market to drive changes. This is a fascinating idea if it can be done. Here's how it works. Sornette and Filiminov model price time series as being generated by a statistical "point process" -- the idea is to generate price dynamics by modelling the arrival of actual buy and sell orders in the market. The simplest way to do this is to use a Poisson process, with equal probability at all times. This gives a random time series of price movements, but an unrealistic one that lacks the most interesting properties of real markets -- fat tails in the distribution of returns, and long term memory in the volatility (and also volume fluctuations). To get realistic time series, it's possible to let the arrival of buy and sell orders have strong correlations in time, as they in fact do in real markets. This technique is referred to as a "self-excited Hawkes model."

Another way to put this is as follows. In an ordinary Poisson process, the average number of events striking in an interval dt (say, 1 second) is a constant, λ. In the richer process with correlations, this will now be a function of time λ(t). The key to the analysis here is expressing this quantity (essentially, the rate of buy and sell orders hitting to book around time t) as the sum of two very different processes -- 1. a background contribution due to external events such as news, which drive the market, and 2. a feedback contribution coming from the tendency for orders now to have consequences, leading to further orders in the future. The result is eq.(1) of the paper:
Here the first term on the right is the background (which drives the exogenous dynamics) and the second term is the feedback, with h being some function that reflects the likelihood that an event at time ti generates another one at time t later. The first term creates a steady stream of events, the second one creates events which create events which create events, a branching stream of further consequences.


Now, the task of fitting time series generated by such processes to real financial data is more involved and relies on some standard maximum likelihood techniques. The authors also assume for simplicity that the function h has an exponential form (events tend to cause others soon after, and less so with increasing time). The key parameter emerging out of such fits is n, which can be interpreted as the fraction of events of endogenous origin, or in effect, the fraction of market activity due to internal dynamics. The statistical fit also estimates μ, this being the background level of exogenous shocks, which also rises and falls with time. Sornette and Filimonov use data on E-mini futures on a second by second basis over about 12 years to run the analysis, the key results of which come out in the figure below.


The four parts going downward show volume and price, and then the estimated background and the fraction of events caused by internal dynamics, n. The most interesting feature is the general rise in n over the decade showing an increasing influence of internal dynamics, or events which causes further events through internal market mechanisms. In contrast, the background of exogenous shocks -- information driven dynamics -- remains more constant (except with a spike around the time of the Lehman Bros collapse). From this the authors offer a few comments:
The first important observation is that, since 2002, n has been consistently above 0.6 and, since 2007, between 0.7 and 0.8 with spikes at 0.9. These values translate directly into the conclusion that, since 2007, more than 70% of the price moves are endogenous in nature, i.e., are not due to exogenous news but result from positive feedbacks from past price moves. The second remarkable fact is the existence of four market regimes over the period 1998-2010:
(i) In the period from Q1-1998 to Q2-2000, the final run-up of the dotcom bubble is associated with a stationary branching ratio n fluctuation around 0.3.
(ii) From Q3-2000 to Q3-2002, n increases from 0.3 to 0.6. This regime corresponds to the succession of rallies and panics that characterized the aftermath of the burst of the dot-com bubble and an economic recession.
(iii) From Q4-2002 to Q4-2006, one can observe a slow increase of n from 0.6 to 0.7. This period corresponds to the “glorious years” of the twin real-estate bubble, financial product CDO and CDS bubbles, stock market bubble and commodity bubbles.
(iv) After Q1-2007 the branching ratio stabilized between 0.7 and 0.8 corresponding to the start of the problems of the subprime financial crisis (first alert in Feb. 2007), whose aftershocks are still resonating at the time of writing.
It should be emphasized that the analysis here is only sensitive to endogenous dynamics over timescales of only around 10 minutes or less. This stems from some assumptions necessary to deal with the highly non-stationary character of the data, as trading volume has exploded over the decade. Hence, the lower values of n earlier in the decade could reflect the failure of this analysis to detect important endogenous feedbacks operating on longer timescales (in the burst of the dot-com bubble, for example).


Now for what is perhaps the most fascinating thing coming out of this paper -- the idea that this analysis may be able to distinguish big markets movements caused by real news or other fundamental changes from those more akin to bubbles and caused purely by human behaviour (panics and the like) or algorithmic feedbacks. Sornette and Filimonov looked at two specific events, on 27 April and 6 May 2010, where markets moved suddenly and in a dramatic way. The first was caused by S&P downgrading Greece's debt rating, the second is, of course, the Flash Crash. The same analysis using their method shows strikingly different results for these two events:

 
The "branching ratio" here is just the n we've been talking about -- the fraction of market dynamics caused by internal dynamics. The most significant finding is that while the first event of 27 April showed absolutely no change in this value, expected given the apparently clear origin of this event in external information, the 6 May Flash Crash shows a sudden spike in the internal dynamics. Hence, based on this example, it appears that this parameter n acts like a flag, identifying events caused by powerful internal feedbacks. As the authors put it,
The top four panels of Fig. 3 show that the two extreme events of April 27 and May 6, 2010 have similar price drops and volume of transactions. In particular, we find that the volume was multiplied by 4.7 for April 27, 2010 and by 5.3 times for May 6, 2010 in comparison with the 95% quantile of the previous days’ volume. The main difference lies in the trading rates and in the branching ratio. Indeed, the event of April 27, 2010 can be classified according to our calibration of the Hawkes model as a pure exogenous event, since the branching ratio n (fig. 3D1) does not exhibit any statistically significant change compared with previous and later periods. In contrast, for the May 6, 2010 flash crash, one can observe a statistically significant increase of the level of endogeneity n (fig. 3D2). At the peak, n reaches 95% from a previous average level of 72%, which means that, at the peak (14:45 EST), more than 95% of the trading was due to endogenous triggering effects rather than genuine news.
Do you believe it? It sounds plausible to me. I guess the thing that would be good to see is some thorough tests of the method applied to time series for which we know the origin of the dynamics. That is, take something like a chaotic oscillator and drive it with some external noise with a controllable level and see if this method generally gives reliable results in teasing out how much of what happens is driven by the noise, and how much by the internal dynamics. Perhaps this has already been done in some of the papers describing the development of the self-excited Hawkes model. I'll try to check on this.


In any event, I think this is certainly a provocative and interesting new approach to this old question of internal vs external dynamics in markets. And I actually don't find the result surprising at all that n has increased markedly over the past decade. This is precisely what one would expect as trading moves over to algorithms that react to what other algorithms do on a sub second basis.