RSS Amplifier

Blog

Ilya Kashnitsky

Dr. Ilya Kashnitsky is a demographer @ Statistics Denmark.

ikashnitsky.phdRSS feed ↗20 posts

Latest posts

UK monarchs’ longevity against their people: a demographically correct reanalysis

Note End of March is the time when I remember with warm nostalgia the vivid memories of working alongside and learning from Jim Vaupel, who died untimely on 27th March 2022. He was a brilliant demographer and a vital person who radiated love to demography and influenced generations of researchers in finding and shaping their academic paths. Please read more about Jim on our collective memorial…

Can deep football knowledge guarantee betting success? Systematic evaluation of football pundits’ La Liga predictions

As football fans, we constantly hear pundits making bold predictions. But how often do they actually hit the targets? Let’s check the track record of three popular Russian football pundits. Vadim Lukomskiy , Denis Alkhazov , and Vladimir Grabchak air a fantastic La Liga Preview podcast on YouTube . Each week they discuss the upcoming matchday and offer their bets, one per game. Next week, they…

Beyond Fraud: How IHME Distorts Academic Metrics

Recently, a post on LinkedIn highlighted a Google Scholar profile of an apparently just starting PhD researcher who suddenly started accumulating unbelievable counts of citations to his questionably numerous papers. Usually, such profiles are a clear indication of involvement into the worst publication practices, mostly paper mills. The post generated a predictable wave of comments along the lines…

Sidekick projects may be worthy distractions for an early academic

In the last few days, I’ve been thinking a lot about Claus Wilke’s blog post on the essential need of writing and publishing a lot of papers for academic researchers. This is such a beautifully formulated argumentation for the position that may easily feel like a losing one in this debate. Quite often we hear calls claiming that researchers should publish just one paper per year. I always felt…

Rotate the damn plot

Several days ago I saw a post on LinkedIn by Jornt Mandemakers with some very curious results from Gender and Generations Programme surveys. While representing interesting data, it was a perfect example of a way too common academic dataviz fallacy, and I decided to finally write this blog post. Here is the original plot by Jornt. We are going to replicate it and then redo – all in hope to…

[UPD] Zotero hacks: reliably setup unlimited storage for you personal academic library

About this tutorial In summer 2024, Zotero had a major update to version 7 . The update affected some of the setup routines that I outlined ages ago in the Zotero hacks post . The recipe laid out in that old post helped me painlessly update, move, and maintain my Zotero library for more than a decade; and judging by feedback it did so for many dozens of my friends, colleagues, and just occasional…

#30DayMapChallenge my 25/30 contributions

For four years I’ve been following #30DayMapCallenge with admiration but not daring to commit to it. Producing maps was always an accompanying step in my main activities, and towards the end of a year it never felt possible to focus on producing them daily. This year I decided to cheat and re-publish many of my maps that accumulated over the years, and only produce new ones for a handful of…

Improve your maps in one line of code changing map projections

Did you ever think why we (okay, I’m clearly biased, maybe just many of us, humans) love maps so much? Why do they often work so much better than other types of dataviz? I think 1 what makes the maps work is the speed with which we can recognize familiar shapes, most often countries. That’s why it’s so annoying when these shapes get distorted – it hinders the smoothness of reading the map and…

Geocode address text strings using tidygeocoder

Deriving coordinates from a string of text that represents a physical location on Earth is a common geo data processing task. A usual use case would be an address question in a survey. There is a way to automate queries to a special GIS service so that it takes a text string as an input and returns the geographic coordinates. This used to be quite a challenging task since it required obtaining an…

Save space in faceted plots

Faceting 1 is probably the most distinctive feature that defined the early success and wide adoption of ggplot2 . Small-multiples are often a great dataviz choice. 2 But one common problem is when your panels for the subsets of data requite vastly different amount of space. By default the panels in faceted ggplots are all of the same size. If the data subsets are very different is size – a common…

Easily re-using self-written functions: the power of gist + code snippet duo

Quite often data processing or analysis needs bring us to write own functions. Sometimes these self-defined functions are only meaningful and useful within a certain workflow or even a certain script. But other self-written functions may be more generic and reusable in other circumstances. For example, one may want to have a version of ggsave() that always enforces bg = 'snow' , or a theme_own()…

The easiest way to radically improve map aesthetics

Since R community developed brilliant tools to deal with spatial data, producing maps is no longer the privilege of a narrow group of people with very specific almost esoteric knowledge, skillset, and often super expensive software. With #rspatial packages, maps (at least the relatively simple ones) became just another type of dataviz. Just a few lines of code can reveal the eye-catching and…

Were there too many unlikely results at the FIFA World Cup 2022 in Qatar?

FIFA World Cup 2022 in Qatar saw many surprising results. In fact, too many – some would argue. From the unbelievable loss of Argentina to Saudi Arabia at the very beginning of the group stage, via the loss of the magnificent Brazil to Cameroon at the end of the group stage, to the groundbreaking performance of Morocco who were competitive playing against all the usual grands. Somewhere towards…

What is life expectancy? And, even more important, what it isn’t

It really is a remarkable achievement and maybe a lot of luck that the world mundanely operates with such a complex indicator as life expectancy. Unlike many statistics and quantities of general use that are being monitored and reported regularly, life expectancy is not observed directly. It’s an output of a mathematical model called life table . And as any model it comes with a certain load of…

Show all data in the background of your faceted ggplot

One of the game-changing features of ggplot2 was the ease with which one can explore the dimensions of the data using small multiples . 1 There is a small trick that I was to share today – put all the data in background of every panel. This can considerably improve comparability of the data across the dimension which splits the dataset into the subsets for the small multiples. Better to show right…

Dotplot – the single most useful yet largely neglected dataviz type

Important I have to confess that the core message of this post is not really a fresh saying. But if I was given a chance to deliver one dataviz advise to every (ha-ha-ha) listening mind, I’d choose this: forget multi-category bar plots and use dotplots instead . I was converted several years ago after reading this brilliant post . Around the same time, as I figured out later, demographer Griffith…

Zotero hacks: unlimited synced storage and its smooth use with rmarkdown

Note About this tutorial Here is a bit refreshed translation of my 2015 blog post , initially published on Russian blog platform habr.com . The post shows how to organize a personal academic library of unlimited size for free. This is a funny case of a self written manual which I came back to multiple times myself and many many more times referred my friends to it, even non-Russian speakers who…

See you in Barcelona this summer

Have you been feeling lately that you are missing out the coolest skill-set in academia? Here is you chance to cut in and dive into R. In July BaRacelona Summer School of Demography welcomes dedicated scholars, aspiring or established, to help them migrate to the world of new oppoRtunities. The school consists of 4 modules. You can take them all or choose specific ones. The first, instructed by…

Compare population age structures of Europe NUTS-3 regions and the US counties using ternary color-coding

On 28 November 2018 I presented a poster at Dutch Demography Day in Utrecht. Here it is: The poster compares population age structures, represented as ternary compositions in three broad age groups, of European NUTS-3 regions and the United States counties. I used ternary color-coding, a dataviz approach that Jonas Schöley and me recently brought to R in tricolore package. In these maps, each…

sjrdata: all SCImago Journal & Country Rank data, ready for R

SCImago Journal & Country Rank provides valuable estimates of academic journals’ prestige. The data is freely available at the project website and is distributed for deeper analysis in forms of .csv and .xlsx files. I downloaded all the files and pooled them together, ready to be used in R. Basically, all the package gives you three easily accessible data frames: sjr_journals (Journal Rank),…