11 April 2020

All models are wrong, but some are useful

Note: Also available as a PDF.

Radoslav Harman, my former statistics professor (and also one of my favorite people), is working on a stochastic model of COVID-19 in Slovakia. Below is a picture from my simulation of the model. Red/orange line is the actual number of new positive cases each day – the real data. Blue dots represent outcomes of the simulation, the darker the more frequent. The picture is made of out 200 runs for one particular setting of free parameters b0, α and prefix (their meaning and exact values aren’t important for now).

Simulation of daily new COVID-19 cases

On March 21, 41 people tested positive out of 440. The previous day (March 20), 13 out of 367 tested positive. With only only 20% more tests on March 21 we got 3 times more (41 versus 13) positive results! The model and my simulation predicted about 10 to 30 positive tests from March 20 to March 22. The simulation is clearly off around this time – it overshoots on March 20 and 22 and undershoots on March 21. How could we fix it? My suspicion is that the population tested on March 21 was qualitatively different, for example it could have contained a big group of people returning from abroad. Actually, this is exactly what happened on April 5, when a group of 35 people returning from Austria tested positive. April 5 is the second to last day in the graph and for this day the model underestimated the number of positive new cases by a huge margin.

The result of the model could be improved by providing information about where each tested person traveled and who they interacted with in the last few days (and this is exactly the type of information that needs to be collected for life to get back to normal).
All models are wrong, but some are useful. – George Box
The model is wrong about daily new cases, but actually being right about daily new cases is not its goal. It’s supposed to describe the behavior of the epidemic in the long term. The following graph is similar to the previous one, but contains cumulative/total cases instead. Also, it contains two more days. Again, the red line is the reality while the simulation is blue.

Simulation of total COVID-19 cases
It looks like this configuration of parameters b0, α and prefix fits the cumulative data fairly well, since the red line stays within the simulated blue zone (except for the last two days when we know that the tested population was somewhat special). As Rado said himself, right now the model admits two scenarios fairly different from each other: one where the disease is spreading slowly since mid-February and one with the disease spreading since the end of February, but much more rapidly.

Is such a model useful? Sure, it could be more precise, but at least it gives a range and that’s a start. I’d say that it’s fairly useful. And as always, we need more and better data.

Exponential-growth models no longer useful

Exponential growth has been mentioned a lot regarding epidemics and the spread of COVID-19. However, in countries implementing social distancing, exponential-growth models are so far from reality that they are no longer useful. And compared to the model above, they cannot be improved using more and better data.

Exponential functions are straight lines on semi-log graphs, so the total number of cases should form a straight line on a semi-log graph. Below is a picture of total cases in Italy on a semi-log graph together with the exponential function 220 ⋅ 1.145t, where t is the number of days since the 200th case.

Total confirmed cases in Italy on a semi-log graph.
The number of cases is not a straight line at all! Clearly, we haven’t applied enough logarithms, so let’s add one more log to the mix and plot it on a log-log graph. Note that now also the horizontal axis is log-scale.

Total confirmed cases in Italy on a log-log graph.
Now it looks much more like a straight line! For about 20 out of the 40 days plotted, it fits the function 2.5 ⋅ t3 which is the straight dashed line on the graph. That’s a polynomial and not an exponential function. Is it just a coincidence?

According to the paper Fractal kinetics of COVID-19 pandemic by Ziff and Ziff published in February 2020, the growth of active cases in China was much slower than exponential – roughly 0.0854 ⋅ t3.09 ⋅ e − t/8.9. Notice the exponent of 3.09 being almost the same as 3 for Italy, though the formula plotted above for Italy didn’t have exponential decay. There is another important difference – the two graphs for Italy contain total confirmed cases while Ziff and Ziff quantify active cases. In particular, active cases can decrease while total confirmed cases never decrease, since it’s the number of people who tested positive.

A few weeks ago, Slovak mathematicians Katarína Boďová and Richard Kollár looked at multiple countries using the same principles presented by Ziff and Ziff, and made a generalized formula for the number of active cases N(t) on day t (t = 1 is chosen as the first day with 200 active cases).

N(t) = (A/TG) ⋅ (t/TG)6.23 / et/TG
  • TG is a country-specific parameter – the average speed in days in which “the sick are removed from the system” (to stop spreading the disease).
  • A is a scaling constant.
  • 6.23 is not just some constant that fits the data but a root of an important equation according to the authors. We’ll have to wait for the manuscript to know more details.
On 2020-03-30, they made the following predictions for active cases in 7 countries.
Country TG A Max. cases Max. date
USA 10.2 72329 1 241 389 2020-05-08
Spain 6.4 3665 99 978 2020-04-12
Italy 7.8 4417 99 459 2020-04-12
Germany 6.7 3773 98 038 2020-04-14
United Kingdom 7.2 2719 66 082 2020-04-21
France 6.5 1961 53 060 2020-04-12
Iran 8.7 2569 51 773 2020-04-21

Scoring the predictions

On 2020-04-10, 11 days after the predictions, I’ve looked at the data to see how did the predictions do. However, I’ve disqualified France, since they first withheld information for multiple days and then reported all of it on 2020-04-04, making it unfair to anyone predicting trends. Slate Star Codex questions Iran’s data, but I haven’t investigated that, so I keep my rating for now.
Country Subjective rating
Italy ★★★★★
United Kingdom ★★★★★
Spain ★★★★★
Germany ★★★★
USA ★★★★
Iran ★★

Italy

The prediction (dashed line) is spot-on. Note that the green zone marks the data available at the day of prediction, so until March 29.
Active cases in Italy until 2020-04-09 together with the predicted trend

Spain

The original prediction (upper dashed line) was a bit pessimistic, but another curve with a slightly lower TG of 6.2 (versus the original 6.4) fits the data really well so far.

Active cases in Spain until 2020-04-09 together with the predicted trends

Germany

Germany, like Spain, at first looked like having lower TG of 6.3, but recently their patients are recovering really fast and the number of new cases is steady. It might mean that their TG got even lower but it could also mean that there is something wrong about the model.

Active cases in Germany until 2020-04-09 together with the predicted trends

USA

USA had the most pessimistic prediction and it’s good that the number of cases is smaller than predicted. I scored it highly (★★★★), because keeping the predicted TG but lowering the value of A by 15% is in line with the trend. However, USA is still before the first inflection point, so it’s too early to make any confident judgements and I’m not confident about my exact score either.
Active cases in USA until 2020-04-09 together with the predicted trend

For the sake of brevity I skipped the two remaining countries, but you can check their predictions on a web dashboard. The data is updated daily.

Discussion

The classic exponential-growth models have a key assumption that infected and uninfected people are randomly mixing: every day, you go to the train station or grocery store where you happily exchange germs with other random people. This assumption is not true now that most countries implemented control measures such as social distancing or contact tracing with quarantine.

You might have heard of the term six degrees of separation, that any two people in the world are connected to each other via at most 6 social connections. In a highly connected world, germs need also very short human-to-human transmission chains before infecting a high proportion of the population. The average length of transmission chains is inversely proportional to the parameter R0 (which you probably heard of).

When strict measures are implemented, the random mixing of infected with uninfected crucial for exponential growth is almost non-existent. For example with social distancing, the average length of human-to-human transmission chains needed to infect high proportion of the population is now orders of magnitude bigger. It seems like the value of R0 is decreasing rapidly with time, since you are meeting the same people over and over instead of random strangers. The few social contacts are most likely the ones who infected you, so there’s almost no one new that you can infect. Similarly for contact tracing and quarantine – it’s really hard to meet an infected person when these are quickly quarantined.

The formula N(t) = (A/TG) ⋅ (t/TG)6.23 ⋅ e − t/TG has two free parameters A and TG. One objection to the finding might be that with two parameters you can create many different functions, making fitting arbitrary curves easy. However, a simple analysis shows that A and TG only scale the graph of the function vertically and horizontally. The observation is left as an exercise to the reader. As a hint, here’s a picture of three functions t6.23/et, (t/2)6.23/et/2 and 2t6.23/et.


Further reading


Back to the first model

The model of Radoslav Harman is an explanatory model: it tries to explain the past. It estimates when the infection arrived to Slovakia and how fast it’s spreading since then. Also, it’s estimating how good the testing selection process is, whether we test too many people with common cold or flu instead of COVID-19 patients.

These three estimation goals are expressed as the following parameters of the model.
  • b0 is a measure of overall quality of the testing selection. That is, the greater b0, the better the efficiency of the system in restricting non-COVID-19 individual from testing. The smaller b0, the more non-COVID-19 individuals get tested.
  • α or γ2 describe the growth of infected cases. The model works both for polynomial growth (parameter α) and for exponential growth (parameter γ2).
    • α is the exponent in the polynomial growth function tα. It’s very hard to estimate TG for Slovakia, since we have few dead and recovered, so there is no support for exponential decay yet.
    • γ2 is the rate of exponential growth after March 12, when Slovakia implemented restrictive policies. Judging by the mobility report released by Google, March 12 seems to be the right choice.
  • tmax is the total length of the simulation. (My code instead uses prefix_length as a parameter and calculates tmax = prefix_length + days, where days is the number of days for which we have data.)
For each combination of parameters, we run multiple simulations. Each configuration of parameters receives a score based on how close it is to reality.
The code is on github, including visualizations and dashboards running on a web server. If you’d like to help in any way, feel free to get in touch.

Picture time!

Let me end with some pictures from the simulations. They confirm Rado’s claim that we cannot distinguish between start of the infection in mid-February versus end of February with faster spread. You can also view the visualizations online: polynomial and exponential growth.

The axes of the heat maps on the pictures show b0 and prefix length. The closer the color on the heat map is to yellow, the better a particular configuration matches reality. Notice that yellow color spans through the pictures, so we have plenty of fairly good configurations. There are two pictures for polynomial growth of t1.24 and t1.28, and one for exponential growth of 1.04t. So far, the scores of these three setups aren’t that different from each other, but we expect it to change in the coming weeks.

Heat map of errors for different b0 and prefix length, with growth of t1.24

Heat map of errors for different b0 and prefix length, with growth of t1.28

Heat map of errors for different b0 and prefix length, with growth of 1.04t

15 February 2020

Stávkovanie pred voľbami

Už platí moratórium na prieskumy, ale stávkovať na výsledky volieb sa paradoxne stále dá. Podľa Efficient-market hypothesis (EMH) by finančné trhy mali reflektovať všetky dostupné informácie. Stávkovanie je tiež finančný trh, čiže v stávkových kurzoch sú už započítané výsledky volebných prieskumov, ale aj ďalšie informácie z “ulice”. Neviem, ako veľmi sa slovenské stávkové spoločnosti približujú ideálu EMH[1], ale predvolebné udalosti určite ešte zahýbu kurzami. Stávkové kurzy sú tak čiastočnou náhradou za chýbajúce prieskumy v posledných 14 dňoch. Ja ich budem brať určite do úvahy pri rozhodovaní, komu dám hlas.

Kurzy tesne pred začiatkom moratória na prieskumy

Zároveň sa dá stávkovanie použiť na emocionálne hedžovanie. Stavil som si na to, že SNS sa dostane do parlamentu, lebo si to neželám.
  1. Ak sa SNS nedostane do parlamentu, znamená to oveľa menšiu šancu pre zlý vládny variant, čo vynahrádza stratené peniaze.
  2. Ak sa SNS dostane do parlamentu, asi bude zlá vláda, ale aspoň som vyhral peniaze. Preto to nazývam emocionálny hedž.
V prípade výhry za polovicu kúpim uhlíkové kredity cez carbonfund.org a za druhú polovicu prispejem jednej slovenskej politickej neziskovke. Napríklad slovensko.digital, ale môžete dať aj iné návrhy.



[1] Slovensko je malá krajina a tak sa volebnými kurzami dá manipulovať pomerne ľahko a ešte ľahšie v lokálnych voľbách. Takto o župných voľbách v Banskej Bystrici písalo sme.sk:
Samozrejme, sú aj situácie, keď preváži stávkovanie srdcom. Napríklad pred poslednými župnými voľbami, kde bol Ján Lunter jasným favoritom so širokou podporou, fanúšikovia Mariana Kotlebu sa na sociálnych sieťach hecovali stávkovaním na Kotlebu a dokonca z neho na chvíľu urobili (kurzového) favorita.

18 June 2019

Kayaking in Möjareservatet

After moving out of Sweden in 2015, I've missed the Stockholm archipelago a lot. There’s nothing quite like it in Europe outside of Scandinavia and British Isles.

Me and Chillu went to my favorite part around the island of Möja. In the beginning of June there were very few people out there, so it was very quiet. As usual, Stockholm archipelago offered the best accommodation for the lowest price of 0 SEK.

The best accommodation is often free

Sunset seen from my sleeping bag

After sunset, as seen from Chillu's sleeping bag
This was our route on the first day (the time and elevation shown by Strava are wrong though).


Slalom around the small islands avoids the open sea, which makes it less windy and more pleasant. Also, the small islands and channels between them are more beautiful than the open sea.

Calm waters in Bockösundet, a channel between two islands

Enough with the words, for more photos you can go to Flickr or Google Photos (a bigger album from two cameras).

11 April 2019

Review of The Bridge from Barbell Medicine

The Bridge v1.0 is an 8-week barbell program to be run once your novice linear progression program stops working. Most of my lifts started stagnating after about 2 months of Phrak’s Greyskull LP (see also my previous post), so The Bridge looked like a good next step.

I chose a program from Barbell Medicine for the following reasons.
The program is organized into lowmoderate and high stress weeks. The first half of the program has higher volume with lower weights, while the second half uses heavy weights and shorter sets, including very heavy singles.

Total weight lifted (in kg) by week for the 4 main lifts

There are 3 strength training days each week complemented by 1 or 2 conditioning days with 30-minute low intensity cardio, 12-minute high intensity interval training, etc.

My results

It took me about 8 and half weeks to finish the 8-week program due to a ski touring trip and an orienteering race. I didn’t have the equipment to do pin squat or pin bench, so I just changed those to the paused versions. Even Austin approves, so I think I followed the program very well.

Before The Bridge, most of my lifts stagnated because of a very weak lower back. I tried to specifically address it, but didn’t see much improvement. On The Bridge, my back got stronger and all the lifts started moving up again. I attribute it to the deadlift variations (rack pulls and paused deadlifts) as well as the tempo squat. Not to mention the magic stress dosage of the program.

e1RM is estimated 1 repetition maximum. The weight is estimated from the number of repetitions and my own RPE rating. Weights are in kilograms.

My weight also went up by about 2 kilograms, which rounds nicely to 10 kilograms gained in 6 months. I can't wear any pants that I wore last year anymore.

I’m really happy with my squat and the increased strength showed on my last ski touring trip. It seems that squat 1RM over 100 kilograms is needed for optimal freeride experience. However, my back is still fairly weak, so my next goal is to improve the deadlift and keep it in the 130 – 150 kilogram range until I’m 89 years old.

My tips

The Bridge is the only program given out for free by Barbell Medicine, which means it’s much less polished than the paid programs, so I recommend you do the following.
  • In case something is not clear, search the forums or r/BarbellMedicine.
  • Use a spreadsheet for tracking your progress and getting weight suggestions for each session.
  • Don’t sweat too much about misjudging RPE. Mike Tuchscherer, who introduced the concept of RPE for lifting (there already was Borg RPE), says not to stress too much about it and even for him it’s hard to distinguish RPE 6 and 7.
  • You can always adjust the weight during the workout, e.g. one day I was supposed to do 3x4@8 and started with 4x85@7.5 (slightly easier than 8) and then continued with 2x4x87.5@8.5 (slightly harder than RPE 8).
  • People complain about spending 2 hours in the gym on the program, but I never took more than 65 minutes when I was alone. I tried finishing one set in 2 and a half minutes. It might have caused lower weight on the bar but it’s a good strategy in the long run.

Conclusion

The Bridge is overall a great program. I might run it again in the future, but probably the upgraded non-free Bridge v3.0.

28 February 2019

A big update to my investing tutorial

I’ve updated my investing tutorial. A few things have changed since I wrote it in 2016.
  • Cheap global currency hedged bond ETFs were introduced by iShares, SPDR and others. TER is only 0.1%, making them some of the cheapest bond ETFs on the market.
  • Cheaper stock ETFs were introduced, mostly by iShares.
  • Buying US-domiciled ETFs is no longer easy for European investors.

27 January 2019

Review of The 4-Hour Body and Starting Strength

The 4-Hour Body by Tim Ferris

The 4-Hour Body is a book covering many topics around diet and exercise. Tim himself doesn’t recommend reading it from cover to cover, so I read about 2 thirds that were interesting to me.

The diet part is the longest. His recommendation to start the day with a high-protein meal is sound, I often do that myself (5 or 6 eggs). However, his obsession with insulin levels seems unnecessary (see Guyenet and Masterjohn). There’s plenty of people who lost a lot of weight on the diet, but I haven’t tried it, since I have the exact opposite problem.

I’ve tried the Occam’s protocol, which is a training program for maximal muscle gain. I ran it for 2 months and gained about 6 kilos of weight, some of which was fat but most was muscle. My rate of weight gain was lower than what Tim promised, which was most likely caused by the following two factors.
  • I ran a few orienteering races while doing the program, but Tim doesn’t recommend any cardio. This was intentional on my side.
  • I started eating as much protein as Tim suggested, but felt horrible. 7 days into the program, a blood test revealed that my kidneys couldn’t keep up with the protein intake. I immediately started eating less and then slowly increased protein intake, but I’ve started feeling strange “in the kidneys” again. This must have been the first time I was aware of my kidneys.
Part of the weight gain might have been regression to the mean, since people’s weight is typically the lowest at the end of the summer which is when I started the program. Regression to the mean could only explain half of the weight gain, though.

I’m convinced the program works, probably much better if you’re able to tolerate more protein. However, my biggest gripe is how Tim sells it. Tim likes to mention the 80/20 rule: you can get 80% of the results by only doing the important 20% of the effort. The title of the book comes from the 4 hours Tim spent in the gym during the month he gained an enormous amount of muscle mass.

The hard part of Occam’s protocol is eating enough protein and calories, which Tim mentions himself many times. I had to think constantly about food and when I managed to hit the protein target, I felt full, lethargic and overall terrible. It felt like 80% of effort for 80% of results.

When the 2 months ended, I switched to Starting Strength which in my opinion achieves what The 4-Hour Body tries to be. For example, Tim has a separate section on how to fix tight muscles by time-consuming exercises or expensive methods for which he flew to another city, but I managed to fix my tight hamstrings just by doing the exercises in Starting Strength. I got much more bang for the buck by following Starting Strength, but more on that in a separate review.

Over the years, I’ve listened to a couple of Tim’s podcasts and some of his interviews, but I have never become a big fan. After reading his book I think I know why. He's a great salesman, but the things he's selling aren't that good.

Starting Strength by Mark Rippetoe

Starting Strength is a beginner barbell program with only 5 exercises (squat, deadlift, over-head press, bench press and power clean). All exercises are explained in great detail—just the squat chapter has about 60 pages.

I didn’t run the Starting Strength linear progression exactly, but rather went with Phrak’s Greyskull LP Variant recommended on Reddit. The two programs are close enough: 4 out of the 5 exercises are the same and you also do 3 sets of 5 repetitions. Weight increases and deloading are the same too.

Starting Strength focuses on strength gain. However, the best part about doing the program were bonus improvements that I didn’t expect.
  • Improved hamstrings flexibility: I suffered from tight hamstrings whenever I played hockey or rode a bike. I’ve tried massages, foam rolling, hot baths and stretching, always achieving only a short-term effect. Just before starting the program I couldn’t touch my toes and now I can press all my fingers against the ground (no palms yet). And no, stretching is not part of the program, so such a huge progress was very surprising to me.
  • Much less muscle soreness: I used to have too much muscle soreness from my semi-regular sport training, but now that is not an issue anymore. I also don’t remember this being promised anywhere, so that’s another surprise.
  • Muscle/weight gain: Even though it’s primary a strength program, I’ve gained about 1 kilogram per month on the program.
  • More strength everywhere: I used to subconsciously search for things to lean against while standing. These days, standing is just easy. Also, free-ride skiing used to be very hard on my legs, but now it feels at least 3 times easier. This is not that surprising, as the exercises hit almost all the muscles.
In the end Starting Strength fulfills the goals of The 4-Hour Body much better, at least for me. Only 5 exercises packaged in a simple-to-follow program feel like the 20% of effort that bring me 80% of results. Yes, my hamstrings could get more flexible by doing more specific work, but as it is now it’s all good.

Note that Mark Rippetoe can be dogmatic, e.g. he repeatedly says in the book that the squat works your hamstrings a lot, but there’s now good scientific evidence that that’s not true. Starting Strength is a great start but it’s definitely not the end.

Conclusion

I’m now well in my 30s and I think now is the right time to get serious about health. Thanks to a lot of trial and error in the last 5 years, I’ve conquered regular insomnia and migraines. Barbell training is another great addition, especially the squat and deadlift feel like cheat codes to life.

The conventional thinking is that your health deteriorates with age but I’d like to keep improving as long as possible. What if there are more cheat codes to be discovered?

23 June 2018

Poľana – mať stará ohromných stínov

Poľana is an inactive stratovolcano in Slovakia not discovered by tourists. I’ve never been there, so I’ve decided to do a multi-day hike there with Ivan and Roman.

Friday

We got off the bus in Strelníky right at sunset. With our headlamps on, we hiked the steep slope to the shelter Partizán nad Mincou.

Hiking towards the shelter

The shelter was empty, so we had a very luxurious stay. At 4AM, we were woken up by a dormouse climbing the walls above us. We managed to scare it off and went back to sleep.

Saturday

In the beginning of our hike we had good views of the surrounding areas. Then we entered the forest and didn’t have views, but the forest was beautiful on its own.

Poľana forest

We ate lunch at the top of Strunga, with excellent views of the whole caldera. When Ivan read Andrej Sládkovič’s poem Detvan, I had goosebumps.

Poľana caldera seen from Strunga

The title of this post is the second line of the poem: Poľana – old mother of great shadows. Once the goosebumps went away, we all concluded that the poem is basically about sex, just like 90% of literature.

In the late afternoon we finally met the first hikers around the highest peak of the mountain range – Poľana at 1458 m. We slept at útuľňa Javorinka, a recently renovated shelter with awesome views.

Útuľňa Javorinka

Sunset

Sunday

I woke everyone up at 4:30 AM to catch the sunrise. Unfortunately, it was too cloudy. However, the Sun rose at the 55-degree bearing over Kráľova hoľa 55 kilometers away. That was an impressive coincidence.

Sunrise over Kráľová hoľa
We went to bed again for 2 more hours and then descended to Hriňová.

Old shepherd's shelter

Small farms of Hriňová, Slovak Tuscany

We ran out of water and tried to get it from the 3 different water sources, but we either couldn’t find them or they were dry. Luckily, we were saved by the helpful people in Zánemecká. From here on, we continued on the bus or train.

The end

After this trip, I only had one question in my head: Why haven't I been to Poľana sooner? It’s not too busy, has a lot of well-preserved nature and awesome tourist shelters.

Small selection of my photos is on Flickr and a bigger album on Google Photos.