arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:1710.03705v2 [cs.CY] 24 Mar 2019

Analyzing gender inequality through large-scale Facebook advertising data

David Garcia    Yonas Mitike Kassa Affiliation:  Complexity Science Hub Vienna, Vienna, Austria Affiliation:  Medical University of Vienna, Vienna, Austria    Angel Cuevas    Manuel Cebrian Affiliation:  Universidad Carlos III de Madrid, 28911 Leganes, Spain    Esteban Moro    Iyad Rahwan Affiliation:  Department of Mathematics & GISC,Universidad Carlos III de Madrid, 28911 Leganes, Spain    Ruben Cuevas Affiliation:  Universidad Carlos III de Madrid, 28911 Leganes, Spain Affiliation:  The Media Lab, Massachusetts Institute of Technology, Cambridge MA, USA Affiliation:  Institute for Data, Systems & Society,Massachusetts Institute of Technology, Cambridge MA, USA Affiliation:  IMDEA Networks Institute, Madrid, Spain Affiliation:  Data61, CSIRO, Melbourne, Australia
Abstract

Online social media are information resources that can have a transformative power in society. While the Web was envisioned as an equalizing force that allows everyone to access information, the digital divide prevents large amounts of people from being present online. Online social media in particular are prone to gender inequality, an important issue given the link between social media use and employment. Understanding gender inequality in social media is a challenging task due to the necessity of data sources that can provide large-scale measurements across multiple countries. Here we show how the Facebook Gender Divide (FGD), a metric based on aggregated statistics of more than 1.4 Billion users in 217 countries, explains various aspects of worldwide gender inequality. Our analysis shows that the FGD encodes gender equality indices in education, health, and economic opportunity. We find gender differences in network externalities that suggest that using social media has an added value for women. Furthermore, we find that low values of the FGD are associated with increases in economic gender equality. Our results suggest that online social networks, while suffering evident gender imbalance, may lower the barriers that women have to access informational resources and help to narrow the economic gender gap.

The Web was designed to be universally accessible and open, carrying the promise of equal opportunity in the access to online information and services [1] as the great potential equalizer [2]. However, despite the widespread adoption of the Web and other Information Communication Technologies (ICT), online access is heterogeneously distributed across demographic factors, such as income and gender –a phenomenon called the digital divide [3, 4, 5].

Governments and global organizations express their concern about the digital divide, aiming to connect the 4 Billion people that remain offline [6, 7]. However, the effects of increasing Internet penetration in development are rarely backed up against empirical data [8], and the latest report by the World Bank suggests that unequally distributed growth in Internet penetration might exacerbate socio-economic inequalities [7]. Beyond the divide in the access to Internet, there are further challenges with respect to digital inequality: the heterogeneity of online activity and engagement across demographic groups [2].

Among online resources, social media play a key role in economic development, for example by providing information that facilitates finding employment [9, 10]. An important open question is whether equality in the access to social media can work as a digital provide [11], bringing equality in other social, political, and economic aspects of society. The World Wide Web Foundation reports that one of the key elements in the digital divide is gender inequality [12]. Social media data show the traces of gender inequalities, from content biases and activity on Wikipedia [13, 14] to visibility and interaction disparities on Twitter [15, 16, 17] and professional gender gaps in LinkedIn [18]. Empirical analyses of digital traces has the potential to track more general demographic patterns [19, 20], such as fertility rates [21].

To properly understand the digital divide, a pervasive problem in cross-country comparisons is the limited size of country samples and the challenges to generate unbiased survey data [7]. To overcome this issue, we deployed a system to collect large-scale data from the Facebook online social network through its marketing API, as explained more in detail in the Materials and Methods section. The Facebook marketing API has been useful in previous research to estimate the value of user data [22], to approximate the size and integration of migrant populations [23, 24, 25], and to generate estimations of Internet and mobile phone gender gaps that explain 69% of the variance of ITU measurements [20]. For our study of the relationship between social media gender divides and other economic, education, heath, and political gender inequalities, we generated an anonymous dataset with statistics about the total number of registered users and daily active users of each gender in each country. While our dataset does not contain personal information on any individual user, our study covers a total of 217 countries and more than 1.4 Billion users.

Our dataset allows the quantification of Facebook activity ratios of each gender in each country. From them, we calculate the Facebook Gender Divide (FGD) as the logarithm of the ratio between the activity ratios for men and for women (see Methods for more details). The FGD has a value below zero when women tend to be more active on Facebook that men, a value close to zero for equal activity tendencies, and a positive value when men are more active on Facebook than women in a country. Our computation of the FGD is consistent with similar measurements constructed from limited survey samples from the Pew Research Center and the Global Web Index, as we comment in the Methods section and show in Supplementary Text 1.

Furthemore, the Facebook marketing API allows us to make precise estimates of the Facebook penetration in a country, calculated as the total number of user accounts (independent of gender and activity) over the total population of the country. We combine these new measurements with standard socio-economic indices, including Gross Domestic Product (GDP), Internet penetration, and economic inequality, as well as indices from the World Economic Forum Gender Gap Report that measure gender equality in terms of education, health, political participation, and economic opportunities [26].

Refer to caption
Figure 1: The Facebook Gender Divide across 217 countries. Countries are colored according their Facebook Gender Divide (FGD), from highly skewed towards males (red), balanced (blue), and towards females (green, not visible). The left inset shows the scatter plot of male and female activity ratios across all countries, revealing a spread along the diagonal. The right inset shows the histogram of FGD values in bins of width 0.2. While the mode of countries is slightly below zero, there is significant skewness towards high FGD values. An online interactive version of this figure can be found in https://dgarcia-eu.github.io/FacebookGenderDivide/Visualization.html.

Results

Fig. 1 shows a world map with countries colored according to their FGD, revealing that many countries are very close to gender equality in Facebook (blue color). The red scale shows countries with positive FGD–that is, higher proportion of males on Facebook. The range of values towards FGD below zero (more tendency for women to be on Facebook) is much narrower than above zero, as can be seen in the scatter plot with the activity ratios of each gender (Fig. 1, left inset), and in the skewness of the distribution of FGD across countries (Fig. 1, right inset).

Countries with high FGD are located around Africa and South West Asia, as shown on Fig. 1. This suggests that variations in socio-economic factors of gender inequality across regions could be explanatory of the FGD. We test this observation using a linear regression model of the FGD as a function of the four indices of gender equality measured by the World Economic Forum (economic opportunity, education, health, and political participation), plus five non-gender-based controls of Internet penetration, population size, economic inequality, Facebook penetration, and mean Facebook active user age (see Methods). The left panel of Fig. 2 shows the quality of the model fit, comparing empirical values of FGD rank versus model predictions. Remarkably, the model can explain well the ranking of FGD (R2=0.74R^{2}=0.74), with very few points far from the diagonal. While this result might be partially explained by Facebook using vital statistics in their calculations, it is nevertheless consistent with replications of the model using limited survey samples from the Pew Research center and the Global Web Index (see Supplementary Text 2). This indicates that the performance of the model is not an artifact of the Facebook marketing API.

Refer to caption

Figure 2: Regression results of FGD as a function of gender equality. A) Model predictions versus rank of FGD, where rank 1 is the country with the highest FGD. The model achieves a high R2R^{2} above 0.740.74, explaining the majority of the variance of the FGD ranking. Some countries are labeled, from high FGD (Liberia, India, and Saudi Arabia) to low FGD (Finland, Norway, and Uruguay), as well as some outliers (Dominican Republic, Austria, and Sri Lanka). B) Coefficient estimates and 95% CI of the terms of the regression fit (excluding intercept). Education (Edu), health (Heal), and economic gender equality (Eco) are significantly and negatively associated with the FGD, but political gender equality (Pol) is not. From the control variables, Internet penetration (IP) is negatively associated with FGD, but the rest are not. The main role of education equality in FGD can be observed on panel A, where dots are colored according to the rank of education gender equality, showing that countries with low FGD are ranked high on education gender equality. An online interactive version of this figure can be found in https://dgarcia-eu.github.io/FacebookGenderDivide/Visualization.html

The right panel of Fig. 2 shows the estimate of the coefficients of our model of FGD. The strongest coefficient is that of education gender equality, which can also be observed on the colors of the left panel of Fig. 2. Specifically, countries with high rank in this index have, on average, lower FGD. Health and economic gender equality also have significant negative coefficient estimates, showing that the FGD captures more than one type of inequality. Note that the index for political gender equality does not have a significant relationship with FGD when the other indices are considered in the model.

Among gender-independent controls, only Internet penetration is negatively associated with FGD. Nevertheless, the FGD is also correlated with GDP per capita (Spearman correlation 0.57-0.57 , p<106p<10^{-6}). For that reason, we repeated the model using GDP as a control variable, finding similar results. These results evidence that the relationship between gender equality indices and FGD is observable when development metrics are considered. We present these additional controls, regression diagnostics, and robustness tests in Supplementary Text 2, concluding that the negative relationships between FGD and gender equality indices are robust.

The value of being active in social media might vary across genders, which we address in a wide country comparison. The general penetration of a communication channel can increase the value that individuals get for using it, which is an example of a feedback mechanism driven by (positive) network externalities [27], also known as Metcalfe’s law [28]. If there are network externalities on Facebook, the activity ratio of countries should scale superlinearly with the Facebook penetration in each country. This scaling relationship with Facebook penetration might vary for the activity ratios of different genders, which would signal an additional marginal benefit of using Facebook for one gender.

Fig. 3 shows the scaling relationship per gender between the activity ratio and the total Facebook penetration in each country. Lines show the result of a power-law fit between both variables with an intercept and an interaction term for gender. The estimate of the scaling exponent for each gender is clearly above one for both genders, revealing a superlinear trend consistent with network externalities in Facebook. This exponent is significantly stronger for female users (αF=1.45\alpha_{F}=1.45 CI=[1.41,1.49][1.41,1.49]) than for male users (αM=1.20\alpha_{M}=1.20 CI=[1.16,1.24][1.16,1.24], see Supplementary Text 3 for details), suggesting that the network externalities in Facebook are stronger for women than for men.

Figure 3: Gender differences in network externalities on Facebook. Scaling of Facebook activity ratio per gender versus total Facebook penetration. Solid lines show fit results and shaded areas show their 95% confidence intervals. Both male and female activity ratios grow superlinearly with Facebook penetration (α>1\alpha>1), indicating positive network externalities. These network externalities are stronger for female than for male users (αF>αM\alpha_{F}>\alpha_{M}).

Given the network externalities shown above, could the FGD be related to changes in economic gender inequality? We test this possibility by analyzing the change in FGD and Economic gender equality between 2015 and 2016. We fitted two regression models, one of changes of Economic gender equality as a function of FGD (FGD2015ΔEco2016FGD_{2015}\rightarrow\Delta Eco_{2016}), and the converse one (Eco2015ΔFGD2016Eco_{2015}\rightarrow\Delta FGD_{2016}), including controls for autocorrelation and GDP as explained in the Methods section. The coefficient estimates, shown in Fig. 4, reveal a significant positive relationship between the FGD rank and changes in economic gender inequality, but not vice versa: there is no significant relationship between Economic gender equality and the changes in FGD.

The partial R2R^{2} value of FGD2015FGD_{2015} in the first model is much higher than the equivalent of Eco2015Eco_{2015} in the second model (median bootstrap values of 0.0270.027 and 0.0020.002 respectively), as shown in the second column of Fig. 4. This suggest the existence of an association between FGD and changes in Economic gender equality such that countries with a low value of FGD (i.e. high rank number) tend more on average to approach economic gender equality. This observation is consistent across age groups and is robust to the inclusion of further control variables including socio-economic indicators, other gender equality metrics, and Hofstede’s culture values [29] (see Supplementary Text 4 for more details). On the contrary, this association is not observable for education gender inequality, as a model of ΔEdu2016\Delta Edu_{2016} shows no significant coefficient for FGD2015FGD_{2015}.

Figure 4: Analysis of changes in economic gender equality and FGD. Coefficient estimates of the regression model of changes in economic gender equality as a function of FGD and control terms (excluding intercept, top left), and of the model of changes in FGD as a function of economic gender equality and control terms (excluding intercept, bottom left). Right panels show the bootstrap distributions of partial R2R^{2} of FGD2015FGD_{2015} in the first model and of Eco2015Eco_{2015} in the second one, with dashed vertical lines showing the median R2R^{2} values: 0.0270.027 in the first model and 0.0020.002 in the second one. The FGD explains changes in economic gender equality much better than economic gender equality explains changes in the FGD.

Discussion

By quantifying the Facebook Gender Divide among 1.4 Billion Facebook users, we demonstrate a number of phenomena that deserve further investigation. The FGD is associated with other types of gender inequality, including economic, health, and education inequality. While the mechanisms behind this connection and its generalizability to other social media remain as open questions, this work is an example of how publicly accessible social media data can be used to understand an important social phenomenon.

Recent reports warn about the possibility that individual Facebook user data was misused by Cambridge Analytica [30], pointing to general concerns about privacy in social media. We share those concerns, in particular with respect to the use of sensitive data in potential conflict with the EU General Data Protection Regulation [31] and regarding the possible construction of shadow profiles of non-users [32, 33]. Nevertheless, our results shows that non-personal data, e.g. anonymous and aggregated data produced by billions of Facebook users can be used for social good, in particular to understand the issue of gender inequalities in society at large.

We found evidence of gender-dependent network externalities—that is, women might receive higher marginal benefit than men from the general adoption of Facebook in a country. While we only observe traces of this phenomenon at an aggregate level, in the differences between activity rates across countries, these results point towards a new research direction: using observational data to understand the value of social networking sites across demographic attributes.

The FGD provides an inexpensive and accessible way to compute gender divides in social media that can be tracked over time and across the vast majority of countries. This allowed us to identify a relationship between the FGD in 2015 and changes in economic gender inequality in 2016. This relationship could be produced by three mechanisms: i) a causation path between the FGD and changes in economic gender equality, ii) a more complex causation from economic gender equality on changes in FGD, or iii) by the prevalence of a third factor of cultural gender norms that drive both the FGD and economic gender equality. While we find evidence for the first explanation, we must note that the real interplay between the FGD and economic gender equality is probably a combination of all three mechanisms, and only future research with more detailed data can answer how.

Our results show trends across a wide range of countries, but caution should be taken when extrapolating to the future or when predicating about individual countries. Before doing so, we need longitudinal models of changes in development factors in a wide range of countries, to find the role of the FGD in broader development sequences that include economic development, health, education, and inequality [34] before formulating policy suggestions. Nevertheless, our results allow us to speculate that social media can be an equalizing force that counteracts other barriers—e.g. those that limit women’s mobility [35]—by providing access to greater economic opportunities and social capital. In a similar way as mobile phones increased the life quality of fishermen in India [11], social media might work as a digital provide that helps disfavored groups despite the still generalized inequalities in access to ICT and in adoption of social media technologies.

Materials and methods

The Facebook Global Dataset

We collected the number of Facebook users by age and gender in each country using the Facebook marketing API [36]. Among other services, this API delivers data for its commercial customers to provide targeted advertising. When supplied with a specific target population, the API returns the total audience size and the price to reach that target audience through Facebook. We iterated over each combination of age and gender values, retrieving the total number of users and the number of Daily Active Users (DAU) for each segment in each country. Our dataset contains the number of male and female registered users and DAU for all available countries11 1 The API does not deliver data for certain countries, e.g. Syria, Iran, and Cuba.. After removing entries of small countries with missing values or low resolution, our dataset contains the total number of users and DAU segmented by age and gender for 217 countries. Age data in the API starts at 13, increasing by one year up to a last bin that contains all users aged 65 or older. We distribute a dataset to allow the replication and extension of our results through a Github repository22 2 https://github.com/dgarcia-eu/FacebookGenderDivide.

Ethical considerations

Our analysis of data from the Facebook marketing API only includes aggregated public information. Even though the sample includes data from underage Facebook users, we had no access to any personal identifiable information of any user and we did not interact or manipulate any research subject. The data retrieval was performed as part of the TYPES project funded by the European Comission (GA-653449) and was approved by the Committee of Ethics in Research of the Carlos III University of Madrid (Ethics Report CEI-2015-001). Our analysis of the data, in line with the growing consensus in ethics [37], is exempted from ethics review, as agreed by the board of IMDEA Networks and by the executive office of the Complexity Science Hub Vienna. Nevertheless, following the guidelines of the Association of Internet Researchers [38], we consider the possible downstream consequences of our large-scale research. The resolution of the Facebook marketing API prevents the singling out of individual users, which makes all our codes useless to identify individuals of any minority or threatened group. In addition, there is no way to identify the accounts of users and employ our analysis for any kind of personalization or individual manipulation. From the onset, our project had the potential to reveal important relationships between social media use and gender inequalities online and offline. These benefits greatly outweigh the minimum risks of analyzing this kind of aggregated data that is accessible to anyone with an Internet connection.

Validating the FGD

Facebook provides the raw data for our study as aggregated values, but as with any new research method, we should not take it at face value without comparing it with more established methods. This is of special importance given the challenges previously found with health-related data from this API [39].

To validate our measurements, we use three reference survey datasets: the Global and Internet & Technology surveys of the Pew Research Center 33 3 http://www.pewglobal.org/dataset/spring-2016-survey-data/44 4 http://www.pewinternet.org/dataset/march-2016-libraries/ and the survey of the Global Web Index (GWI)55 5 https://www.globalwebindex.net/. These datasets allow us to compute reference measurements of Facebook penetration and FGD for small samples of countries, to be compared with our calculation of the FGD through the Facebook marketing API.

The results of this validation exercise are reported in detail in Supplementary Text 1. We find high correlation coefficients between our measurement of penetration and the equivalents in GWI and the Pew Global survey, and for the case of FGD as well. These correlations are as good as the correlations between survey datasets, showing that the Facebook API data has comparable quality but a much higher coverage in terms of countries and better temporal resolution. We find low and non-significant correlations between the absolute difference between our measurement of FGD and the one from surveys, but nevertheless add a control for Facebook penetration in our models to make sure that our results are not an artifact of correlated errors in the quantification of FGD.

We further compare Facebook penetration across ages groups in the US through the Pew Internet & Technology survey and the GWI. We find very high correlations between age-dependent measurements. In addition, we explore how representative is the FGD for gender divides in other social media, as captured by the GWI survey. We found moderate yet significant correlations with other media such as WhatsApp, Twitter, and YouTube. This shows that, while we should not take Facebook as representative for all social media, there is certain similarity in gender differences that can motivate future research.

Finally, we test for intra-day oscillations of the measurement of FGD and Facebook penetration and found extremely consistent values. For the case of the FGD and the network externalities model, we also repeat our analysis on monthly snapshots of Facebook data for a period of twelve months between 2015 and 2016, calculating median DAU values each month. This way, we can confirm the robustness of our analysis to possible temporal changes in the way Facebook reports data through their API.

Gender equality and development datasets

To normalize the number of active users over the total population of each country, we use the data collected by the US Census Bureau International Database66 6 https://www.census.gov/programs-surveys/international-programs/about/idb.html. This dataset contains estimates of the resident population by age and gender for more than 226 countries. We combine this data with gender equality indices measured by the World Economic Forum Gender Gap reports of 2015 and 2016 [26]. This dataset quantifies the magnitude of gender equality in 145 countries, measuring it with respect to four key areas: health, education, economic opportunity, and politics. This report updates the values for education, economic, and political gender equality on a yearly basis, allowing us to measure changes between 2015 and 2016. To account for additional economic and development indicators, we include data from the World Bank and the Human Development Index [40], measuring control variables of GDP (PPP) per capita in 2012, economic inequality as the quintile ratio, and Internet penetration.

Computing the Facebook Gender Divide

We quantify the Facebook Gender Divide as a comparison of the rates of activity between genders. The DAU measures how many users have logged into Facebook at a given day, which could be either through a Web browser or a mobile application. We use the segmented data from 13 to 65 years old to normalize the DAU over the total population of a country in those ages, truncating all data that is not included in that age range. This way we avoid introducing a bias with life expectancy and average age. To have a stable estimate of the DAU, we use as the median value over the month of July 2015, replicating over other months afterwards. This way, for each country cc and gender g[Female,Male]g\in[Female,Male]77 7 For simplicity, we take gender as birth sex, i. e. male or female., we have a measurement of the number of active users Ag,cA_{g,c} between 13 and 65 years old. Additionally, this allows us to calculate the mean user age for a country to include it in our models.

Using the US Census Bureau data, we calculate the total population of each gender between the ages of 13 and 65 years old in each country, which we denote as Pg,cP_{g,c}. This way, we can normalize the total activity in Facebook over the population in the same age ranges, calculating the activity ratios Rg,c=Ag,c/Pg,cR_{g,c}=A_{g,c}/P_{g,c}. We define the Facebook Gender Divide in country cc as

FGDc=log(RMale,cRFemale,c)FGD_{c}=log\left(\frac{R_{Male,c}}{R_{Female,c}}\right)

, which compares male and female Facebook activity rates over the population of country cc. A country with positive FGD will have a tendency for men to be more present on Facebook, while a country with negative FGD will show the opposite tendency. A country with FGD=0FGD=0 will have complete equality in the activity tendencies of both genders.

We further compute the Facebook penetration as the ratio between user accounts between 13 and 65 years old reported by the API (regardless of activity and gender) and the total population of the country between those ages.

Regression Models

We model dependencies between gender equality indicators and the FGD as linear models, after applying a rank transformation to all variables such that rank 1 is the highest possible value of the variable. This way, we explore monotonic dependencies that do not need to be linear. We define this FGD model as:

FGD=afQ+bfC+cf+ϵFGD=a_{f}\cdot Q+b_{f}\cdot C+c_{f}+\epsilon

where QQ is a matrix with the ranks of economic, health, education, and political gender equality in each country and CC contains control variables such as Internet penetration (IP), income inequality (Ineq), total population (Pop), Facebook penetration (FBP) and mean user age (Age). cfc_{f} is the intercept and ϵ\epsilon denotes the residuals as the normally distributed, uncorrelated error of the model.

We analyze the relationship between changes and levels in economic gender equality and of FGD through with two models. First, an equality changes model:

ΔEco2016=aoEco2015+boFGD2015+coO+do+ψo\Delta Eco_{2016}=a_{o}\cdot Eco_{2015}+b_{o}\cdot FGD_{2015}+c_{o}\cdot O+d_{o}+\psi_{o}

And second, a FGD changes model:

ΔFGD2016=aqFGD2015+bqEco2015+cqO+dq+ψq\Delta FGD_{2016}=a_{q}\cdot FGD_{2015}+b_{q}\cdot Eco_{2015}+c_{q}\cdot O+d_{q}+\psi_{q}

where ΔEco2016\Delta Eco_{2016} and ΔFGD2016\Delta FGD_{2016} is the change in economic gender inequality and FGD between 2015 and 2016. Both models include a control for autocorrelation as a term with the unranked value of the variable in the 2015, and a main term of the rescaled ranked value of the other variable. Following previous Economics research on Facebook data [10], we include various ranked controls in the matrix OO, first with a simple correction for GDP, but then with extensions with other controls as for the FGD model.

We report the coefficient estimates of robust MM regressors for both models. To compare the effects of one variable over the changes in the other, we first residualize the changes by fitting against all controls. Then we compute the partial R2R^{2} value of the conditioning variable when fitting the residualized values. To understand the uncertainty of this analysis, we bootstrap over 10.00010.000 samples and report the distribution of R2R^{2} values.

We model network externalities as a power-law relationship between the activity ratio of a gender (Rg,cR_{g,c}) and the total Facebook penetration for both genders together (PcP_{c}) in a joint model that includes an intercept for gender and interaction with gender. We define this way the network externalities model as:

log(Rg,c)=αlog(Pc)+β+δg,Female(αFlog(Pc)+βF)+ϕlog(R_{g,c})=\alpha\cdot log(P_{c})+\beta+\delta_{g,Female}(\alpha_{F}\cdot log(P_{c})+\beta_{F})+\phi

where α\alpha measures the scaling relationship between the Facebook presence ratio and the activity ratio of male users, αF\alpha_{F} the difference in that relationship for female users, and ϕ\phi the residuals. The Kronecker delta function δg,Female\delta_{g,Female} takes value 11 when g=Femaleg=Female and 00 otherwise.

All the above models do not show relevant multicollinearity when measuring Variance Inflation Factors [41].

We report the fit the FGD model and the network externalities model with Markov Chain Monte Carlo sampling in JAGS [42]. We also fit all models with robust regression [43], reporting the results of the changes models in the main text and of the rest in the Supplementary Information.

To test the validity of the assumptions of our models after fitting, we verify the normality of residuals through Shapiro-Wilk tests [44], and check that residuals are uncorrelated with fitted values and independent variables. For the case of the network externalities model, we additionally analyze multiplicative residuals in order to test for the possible role of outliers, as shown more in detail in Supplementary Information.

Acknowledments

D.G. acknowledges funding from the Vienna Science and Technology Fund through the Vienna Research Group Grant “Emotional Well-Being in the Digital Society” (VRG16-005). E.M. acknowledges funding from Ministerio de Economía y Competividad (Spain) through projects FIS2013-47532-C3-3-P and FIS2016-78904-C3-3-P. A.C. acknowledges funding from the H2020 EU project TYPES (grant no. 653449) and the Ramón y Cajal grant (RyC-2015-17732). Y.M.K. acknowledges funding from the H2020 EU project TYPES (grant no. 653449). R.C. acknowledges funding from H2020 EU project ReCRED (grant no. 653417). We thank the Pew Research Center and the Global Web Index for the access to their data to validate the FGD.

Supplementary Text 1 - Validating the FGD

We validate the measurement of the FGD against three reference datasets:

  1. 1.

    Global Web Index (GWI). We use the survey responses of the Global Web Index 88 8 https://www.globalwebindex.net/ panel during the period of our study (the two last quarters of 2015 and the two first quarters of 2016). For this period, the GWI contains responses from 99,338 panelists in 34 countries, providing rescaled estimates of survey responses that generalize to the population as a whole. We take the response to the question about the frequency of use of Facebook, calculating the weighted fraction of respondents for each gender and age segment that report to use Facebook daily or more than once a day. Furthermore, we repeat the analysis for other social media (Whatsapp, LinkedIn, Twitter, Instagram, and YouTube), calculating gender divide values outside Facebook.

  2. 2.

    Pew Research Center Spring 2016 Global Attitudes (Pew Global). This dataset includes responses from 23,462 panelists in 19 countries abd can be found online 99 9 http://www.pewglobal.org/dataset/spring-2016-survey-data/. We use the positive answer rate to question 82 “Do you ever use online social networking sites like Facebook, Twitter…?” as way to measure the penetration of SNS in general. We take the respondent weights reported in the dataset to rescale the frequency of positive responses taking into account the self-reported gender of survey respondents.

  3. 3.

    Pew Research Center Internet & Technology, March 7-April 4, 2016 (Pew US). This US questionnaire 1010 10 http://www.pewinternet.org/dataset/march-2016-libraries/ includes questions about Facebook use in particular (act135, “Do you ever use the internet or a mobile app to use Facebook?”), as well as gender and age data from 1,601 respondents. We use the respondent weights to compute Facebook use rates across gender and age groups.

Comparison of Facebook penetration estimates

Figure 5: Total Facebook penetration for both genders as reported by the marketing API versus Facebook penetration as estimated in the GWI survey (left) and penetration of all Social Networking Sites in Pew Global survey (right). The red dashed line shows a linear regression profile, with its prediction standard errors in the shaded area.

The left panel of Fig. 5 shows the relationship between the total Facebook penetration for both genders as measured by us through the marketing API versus the value estimated from the GWI survey. There is a high positive correlation between both measurements of Facebook Penetration (Pearson 0.890.89 , CI [0.78,0.94][0.78,0.94], Spearman 0.860.86, CI [0.72,0.94][0.72,0.94]). The right panel of Fig. 5 shows the same evaluation against the Pew Global survey. The correlation between both measurements is positive and high (Pearson: 0.710.71, CI [0.39,0.88][0.39,0.88], Spearman: 0.770.77, CI [0.43,0.94][0.43,0.94]).

Figure 6: Left: Comparison of GWI Facebook penetration estimate and Pew penetration of all SNS. The red dashed line shows a linear regression profile, with its prediction standard errors in the shaded area. Right: Bootstrapping distributions of Spearman’s correlation coefficient between all pairs of penetration measurements.

The left panel of Fig. 6 shows the comparison between estimates based on GWI and Pew Global survey data. The correlation is also positive and significant, even though samples are of limited size (Pearson: 0.630.63, CI [0.17,0.86][0.17,0.86], Spearman: 0.70.7, CI [0.38,0.91][0.38,0.91]) Both in the right panel of Fig. 5 and in left panel of Fig. 6, it can be seen that China is a clear outlier. This stems from the difference between comparing Facebook penetration versus penetration for Social Networking Sites in general, which is the precise question of the Pew survey. If we focus on the rest of countries, where we can expect a priori that Facebook is more representative of social media in general, the correlation reaches higher values when comparing to the Facebook marketing API (Pearson: 0.840.84, CI [0.62,0.94][0.62,0.94], Spearman: 0.870.87, CI [0.66,0.96][0.66,0.96]) and to the GWI survey data (Pearson 0.840.84, CI [0.55,0.95][0.55,0.95], Spearman 0.780.78, CI [0.43,0.94][0.43,0.94]). To make a fair comparison, we should not include China when comparing Facebook penetration with penetration in SNS in general.

When we focus only on the set of countries available in all three datasets, we can compare the correlation of each pair of data sources to assess the validity of the Facebook API data in terms of Facebook penetration. In this smaller sample, the Facebook penetration calculated through the marketing API is highly correlated with both the Pew Global survey values (Pearson: 0.870.87, CI [0.64,0.96][0.64,0.96], Spearman: 0.860.86, CI [0.59,0.97][0.59,0.97]) and with the GWI survey (Pearson: 0.890.89, CI [0.68,0.96][0.68,0.96], Spearman: 0.830.83, CI [0.53,0.97][0.53,0.97]). The point estimates of these two values are higher than the correlation between the Pew Global and GWI datasets, as reported above. Nevertheless, these differences are not significant, as confidence intervals overlap and bootstrap sampling shows that estimates are indistinguishable (Fig. 5). From this analysis we conclude that the estimate of Facebook penetration from the marketing API has comparable quality to the values reported in the high-quality, representative surveys of GWI and the Pew Research Center.

Facebook gender divide estimates

Figure 7: Measurement of the FGD in the API versus estimates using GWI data (left) and a gender divide for all SNS in PEW (right). The red dashed line shows a linear regression profile, with its prediction standard errors in the shaded area.

We calculated surrogates of the FGD based on the GWI and Pew global survey datasets. The left panel of Fig. 7 shows the comparison of the Facebook marketing API estimate with the same estimate using GWI data, revealing high positive correlation (Pearson: 0.830.83, CI [0.68,0.91][0.68,0.91], Spearman: 0.630.63, CI [0.27,0.87][0.27,0.87]). The right panel of Fig. 7 shows the comparison of the Facebook marketing API estimate with Pew survey data for all SNS (including China), also revealing high positive correlation (Pearson: 0.850.85, CI [0.65,0.94][0.65,0.94], Spearman: 0.740.74, CI [0.35,0.91][0.35,0.91]).

Figure 8: Left: Comparison of gender divide measurements in PEW and GWI data. The red dashed line shows a linear regression profile, with its prediction standard errors in the shaded area. Right: Bootstrapping distributions of Spearman’s correlation coefficient between all pairs of gender divide measurements.

A comparison between estimates of the FGD using GWI versus using Pew data is shown on the left panel of Fig. 8, also revealing positive correlations (Pearson: 0.850.85, CI [0.6,0.95][0.6,0.95], Spearman: 0.490.49, CI [0.11,0.86][-0.11,0.86]). As with Facebook Penetration, the correlation between the FGD using the Facebook marketing API and the other two (with Pew: Pearson: 0.770.77, CI [0.43,0.92][0.43,0.92], Spearman: 0.680.68, CI [0.14,0.93][0.14,0.93]; with GWI: Pearson 0.960.96, CI [0.890.99][0.890.99], Spearman 0.830.83, CI [0.47,0.97][0.47,0.97]) estimates is comparable to the correlation within estimates, as evidenced in bootstrapping samples reported in the right panel of Fig. 8. We can conclude that the estimate of the FGD using the marketing API is consistent with GWI and Pew survey metrics, opening the study of the FGD to a much larger sample of countries.

We measured the absolute difference between the FGD in the marketing API and in each survey dataset. As expected, the correlation between this absolute difference and the Facebook penetration across countries is negative (with GWI: Pearson: 0.3-0.3, CI [0.59,0.05][-0.59,0.05], p-value=0.09=0.09; with Pew: Pearson: 0.27-0.27, CI [0.65,0.21][-0.65,0.21], p-value=0.26=0.26), but its value is weak and not significant. Nevertheless, we include controls for Facebook penetration in our further analyses, to make sure that our results are not an artifact of a correlation between penetration and measurement error in the marketing API.

Furthermore, the GWI survey allows us to compare measurements of the FGD in other social networks with our measure based on the Facebook marketing API. We get moderate to high Pearson correlation coefficients with other sites, such as Whatsapp (0.670.67, CI [0.43,0.82][0.43,0.82]), LinkedIn (0.650.65, CI [0.40,0.81][0.40,0.81]), Twitter (0.690.69, CI [0.46,0.84][0.46,0.84]), Instagram (0.790.79, CI [0.62,0.89][0.62,0.89]), and YouTube (0.890.89, CI [0.79,0.94][0.79,0.94]). While we cannot generalize to all social networks based only on Facebook data, we can see that, to some extent, the difference in activity across genders also appears in other SNS. This is particularly interesting when comparing Facebook, a very private social network, with YouTube or Twitter, which are much more public but still display substantial correlations in terms of FGD.

Comparison across age groups

Figure 9: Comparison of FB presence ratio versus PEW gender categories for both genders together (left) and gender-wise (center). Replication of the same validation versus GWI estimates (right).

We compared Facebook penetration estimates in the US across the four age groups reported in the Pew US dataset. The left panel of Fig. 9 shows the comparison for both genders together, which have a Pearson correlation coefficient of 0.960.96, CI [0.04,0.99][0.04,0.99]. The central panel of Fig. 9 shows the same comparison by taking the gender-wise estimates, which also have positive Pearson correlation (0.940.94, CI [0.72,0.99][0.72,0.99]). This also appears when surveying the GWI dataset for US respondents in similar age categories, as shown on the left panel of Fig. 9, which has a high and significant Pearson correlation coefficient of 0.920.92, CI [0.69,0.98][0.69,0.98]. We can conclude that data provided by the Facebook marketing API is consistent across ages, but to be sure that our further analyses are robust we take two action: 1) we add a mean user age control to our regression models, and 2) we stratify our analyses across age categories, using in each stratum a measurement of the FGD in the corresponding age range.

Subdaily measurement consistency

Figure 10: Examples of subdaily trajectories of the FGD.

We retrieved data from the Facebook marketing API on a daily frequency, starting our retrieval at 3 AM Central European Time. To validate the consistency of our measurement with any other times of the day, we checked the consistency of our construction of the FGD with hourly values for 24 hours in December 2017.

Fig. 10 shows the hourly measurement for a sample of large countries, revealing high consistency with very small fluctuations. When comparing the measurement of the FGD at 3AM CET with any other time in the same day, we get extremely high pearson correlation coefficients (0.99426540.9942654, CI [0.9939277,0.9945844][0.9939277,0.9945844]), as also evidenced in Fig. 11. This also extends to the measurement of Daily Active Users for male (0.99994310.9999431, CI [0.9999397,0.9999462][0.9999397,0.9999462]) and female (0.99993780.9999378, CI [0.9999341,0.9999412][0.9999341,0.9999412]), as well as the total number of accounts for male users (0.99999260.9999926, CI [0.9999921,0.9999930][0.9999921,0.9999930]) and female users (0.99999680.9999968, CI [0.9999966,0.9999970][0.9999966,0.9999970]). Any fluctuation can be attributed to the rounding that Facebook does to preserve individual user anonymity and to the inter day changes in the number of Daily Active Users

Refer to caption
Figure 11: Facebook API measurements at different hours of the day. The left panel shows a comparison of measurements of the FGD for all countries in the dataset at 3AM CET versus hourly measurements at other times of the day. The right panels show the comparison between the measurement at 3AM CET and at other times of the day for the number of DAU and of present users (TOT) per gender.

Supplementary Text 2 - FGD as a function of other inequalities

Regression diagnostics

The Variance Inflation Factors of the variables in the FGD model are below 5, allowing us to discard collinearity in the linear model of FGD as a function of other inequalities. Table 1 reports the detailed results of the FGD model fit and Table 2 reports the results of the same model when fitted with a robust regression method. Table 3 shows a fit with HC correction for heteroskedasticity. All results are qualitatively similar, revealing that the FGD model result is robust to outliers and heteroskedasticity.

Term Median estimate 95% Credible Interval p-value
Intercept 135.8\bf 135.8 [119.9,152.1][119.9,152.1] p<0.01p<0.01
Education Equality Rank 0.54\bf-0.54 [0.67,0.41][-0.67,-0.41] p<0.01p<0.01
Health Equality Rank 0.27\bf-0.27 [0.37,0.17][-0.37,-0.17] p<0.01p<0.01
Economic Equality Rank 0.16\bf-0.16 [0.27,0.06][-0.27,-0.06] p<0.01p<0.01
Political Equality Rank 0.050.05 [0.05,0.14][-0.05,0.14] 0.190.19
Internet Penetration Rank 0.27\bf-0.27 [0.44,0.09][-0.44,-0.09] p<0.01p<0.01
Income Inequality Rank 0.010.01 [0.09,0.10][-0.09,0.10] 0.440.44
Population Rank 0.010.01 [0.09,0.11][-0.09,0.11] 0.400.40
Facebook Penetration Rank 0.030.03 [0.12,0.18][-0.12,0.18] 0.330.33
Mean User Age Rank 0.020.02 [0.08,0.11][-0.08,0.11] 0.330.33
NN 142 R2R^{2} 0.74170.7417
Table 1: Regression results of FGD model. Estimates of p-values are based on the posterior of parameter estimates after 10,000 iterations.

Fig. 12 shows the normal Q-Q plot and the histogram of residuals, which are distributed very close to normality. This is confirmed by a Shapiro-Wilk normality test, with a statistic of 0.99 and unable to reject the null hypothesis that residuals are normally distributed (p=0.63p=0.63). Furthermore, residuals are uncorrelated with all gender equality variables (Near-zero Pearson correlation coefficients, with p-values above 0.90.9) and the square root of absolute residuals are not significantly correlated with predicted values. In addition, Facebook penetration is uncorrelated with residuals of the model (Pearson 0.0028-0.0028, p-value =0.97=0.97), showing no signs of bias due to the variance of Facebook penetration rates across countries.

Term Estimate Standard Error t-value p-value
Intercept 137.4\bf 137.4 9.149.14 15.0415.04 p<1010p<10^{-10}
Education Equality Rank 0.55\bf-0.55 0.070.07 7.47-7.47 p<1010p<10^{-10}
Health Equality Rank 0.29\bf-0.29 0.050.05 4.98-4.98 p<105p<10^{-5}
Economic Equality Rank 0.13\bf-0.13 0.060.06 2.03-2.03 0.040.04
Political Equality Rank 0.050.05 0.050.05 0.950.95 0.340.34
Internet Penetration Rank 0.29\bf-0.29 0.100.10 2.89-2.89 p<0.01p<0.01
Income Inequality Rank 0.0010.001 0.050.05 0.010.01 0.990.99
Population Rank 0.003-0.003 0.050.05 0.050.05 0.960.96
Facebook Penetration Rank 0.050.05 0.090.09 0.620.62 0.540.54
Mean User Age Rank 0.010.01 0.050.05 0.210.21 0.840.84
NN 142 Multiple R2R^{2} 0.67150.6715
Table 2: Robust regression results of FGD model.
Term Estimate 95% HC CI Standard error p-value
Intercept 136.8\bf 136.8 [122.5,151.2][122.5,151.2] 7.327.32 p<1010p<10^{-10}
Education Equality Rank 0.54\bf-0.54 [0.68,0.40][-0.68,-0.40] 0.070.07 p<1010p<10^{-10}
Health Equality Rank 0.27\bf-0.27 [0.36,0.18][-0.36,-0.18] 0.050.05 p<108p<10^{-8}
Economic Equality Rank 0.17\bf-0.17 [0.26,0.07][-0.26,-0.07] 0.050.05 p<0.001p<0.001
Political Equality Rank 0.040.04 [0.04,0.13][-0.04,0.13] 0.040.04 0.320.32
Internet Penetration Rank 0.27\bf-0.27 [0.44,0.10][-0.44,-0.10] 0.090.09 p<0.01p<0.01
Income Inequality Rank 0.0040.004 [0.09,0.10][-0.09,0.10] 0.050.05 0.920.92
Population Rank 0.010.01 [0.08,0.10][-0.08,0.10] 0.050.05 0.870.87
Facebook Penetration Rank 0.030.03 [0.12,0.18][-0.12,0.18] 0.080.08 0.70.7
Mean User Age Rank 0.020.02 [0.07,0.10][-0.07,0.10] 0.040.04 0.660.66
Table 3: Coefficient estimates using HC corrected estimates.
Figure 12: Left: Normal Q-Q plot of residuals of the FGD model. Right: Histogram of residuals.

The role of GDP

Refer to caption
Figure 13: Relationship between FGD and GDP (log scale).

Fig. shows the relationship between the FGD and the GDP per capita. There is a significant negative correlation between both (Spearman 0.57-0.57, p<106p<10^{-6}), motivating a replication of the above model with GDP as a control. Since GDP is highly correlated with Internet Penetration and leads to high Variance Inflation, we replace Internet Penetration with GDP in our model. Coefficient estimates are reported on Fig. 14, revealing that the main result is robust to controlling for the wealth of countries.

Figure 14: Coefficient estimates of the FGD model with GDP instead of Internet Penetration as control.

Model stability across monthly measurements

We repeated the fit of the FGD model for measurements of the DAU in twelve months between 2015 and 2016. Fig. 15 shows the results of the fit for these alternative periods. The coefficient estimates barely depend on the period when the DAU are calculated and the R2R^{2} of the fits range between 0.7270.727 and 0.7580.758, confirming that our results are robust to fluctuations in the reporting of DAU through the Facebook API.

Figure 15: Results of repetition of the fit of the FGD model for 12 months between 2015 and 2016. The results are qualitatively stable across months.

Model stability across age groups

Figure 16 shows the replication of the model for segments of different age groups. Results are qualitatively similar to those of the whole population, with significant effects of gender equality variables and high R2R^{2} values.

Figure 16: Replication of the model for data segmented into age groups.

Model test with GWI and Pew approximations to the FGD

We repeated the model using approximations of the FGD using the limited sample of GWI. The performance of the model is similar, as shown in Figure 17, with R2=0.77R^{2}=0.77. While the sample size of GWI is too small to test the role of all equality variables, the Pearson correlation between the rank of education equality and the rank of FGD in GWI is 0.52-0.52 (pvalue<0.01p-value<0.01). Similarly, the model for the approximation of the FGD with data from Pew for all SNS gives similar R2=0.63R^{2}=0.63 and a significant negative Pearson correlation between the rank of education equality and the rank of FGD in PEW (0.73-0.73, pvalue<0.001p-value<0.001).

Refer to caption
Figure 17: Replication of the model with the GWI estimate of the FGD.

Supplementary Text 3 - Network externalities

The results of the network externalities model are shown on Fig. 18. The model achieves a R2=0.96R^{2}=0.96 on the logarithmic scale and a R2=0.89R^{2}=0.89 on the linear scale of activity ratios per gender. Table 4 shows the detailed results of the model, evidencing the superlinear scaling (α=1.2\alpha=1.2) and the difference between genders (αF=0.25\alpha_{F}=0.25).

Figure 18: Results of network externalities model fit. The left panel shows the comparison between empirical and predicted values, the right panel shows median estimates and 95% CI of model coefficients.
Term Median estimate 95% Credible Interval p-value
β\beta 0.57\bf-0.57 [0.63,0.50][-0.63,-0.50] p<0.01p<0.01
α\alpha 1.198\bf 1.198 [1.16,1.24][1.16,1.24] p<0.01p<0.01
βF\beta_{F} 0.15\bf 0.15 [0.06,0.25][0.06,0.25] p<0.01p<0.01
αF\alpha_{F} 0.25\bf 0.25 [0.2,0.31][0.2,0.31] p<0.01p<0.01
NN 422 (211 countries, 2 genders) R2R^{2} 0.960.96
Table 4: Regression results of network externalities model. Estimates of p-values are based on the null hypothesis that the coefficient equals one for α\alpha and zero for the rest, after 10,000 iterations.

Fig. 19 shows the analysis of the residuals of the model and the error in the linear scale of activity ratios per gender. Some small deviations from normality can be observed at the tails, corresponding to significant Shapiro-Wilk statistics of 0.94 and 0.95. Both types of residuals are uncorrelated with Facebook penetration and do not appear to have a structure across predicted values. We identified some of the residual outliers, such as China and Tajikistan, which when removed do not have a qualitative impact in the results of the model fit and lead to residual distributions closer to normality.

Figure 19: Analysis of residuals of network externalities model. The top panels show the normal Q-Q plot and the histogram of residuals ϕ\phi of the model, and the lower panels the converse for multiplicative residuals in the linear scale of activity ratios per gender. Some minor deviations from normality can be observed in both.

Table 5 shows the results when correcting for heteroskedasticity. All results remain qualitatively unchanged.

Term Estimate 95% Confidence Interval Standard Error p-value
β\beta 0.57\bf-0.57 [0.61,0.53][-0.61,-0.53] 0.020.02 p<1010p<10^{-10}
α\alpha 1.198\bf 1.198 [1.17,1.23][1.17,1.23] 0.020.02 p<1010p<10^{-10}
βF\beta_{F} 0.15\bf 0.15 [0.07,0.24][0.07,0.24] 0.0450.045 p<0.001p<0.001
αF\alpha_{F} 0.25\bf 0.25 [0.18,0.33][0.18,0.33] 0.0380.038 p<1010p<10^{-10}
Table 5: HC corrected results of network externalities model.

We repeated the fit using a robust regression method, reporting the results on Table 6. While estimates slightly change, the qualitative results of a superlinear relationship that is stronger for female users still hold. This shows that our conclusions are robust to the influence of outliers.

Term Estimate Standard Error t-value p-value
β\beta 0.523\bf-0.523 0.0330.033 16.25-16.25 p<1010p<10^{-10}
α\alpha 1.19\bf 1.19 0.020.02 60.8860.88 p<1010p<10^{-10}
βF\beta_{F} 0.25\bf 0.25 0.050.05 5.435.43 p<107p<10^{-7}
αF\alpha_{F} 0.37\bf 0.37 0.030.03 12.5912.59 p<1010p<10^{-10}
Table 6: Robust regression results of the network externalities model.

As with previous models, we evaluated the model of network externalities over twelve months following our initial measurement. Fig. 20 reports the overall results, showing no relevant decrease in R2R^{2} and generally the same result, where the parameter αF\alpha_{F} is significantly larger than zero and the parameter α\alpha is significantly larger than one.

Figure 20: Results of repetition of the fit of the network externalities model for 12 months between 2015 and 2016. The results are qualitatively stable across months.

We stratified the analysis, fitting the network externalities model using calculations of the FGD using only data from a set of age categories. Fig. 21 shows the model results, evidencing that the female intercept, which measures the surplus of the exponent for female users, is positive and significant for all age categories.

Figure 21: Results of repetition of the fit of the network effects model for age segments.

Supplementary Text 4 - Gender equality changes

Table 7 reports the Variance Inflation Factors of the variables in the model of economic gender equality changes and in the model of changes of FGD. All factors are low enough to discard multi-collinearity.

ΔEco\Delta Eco Model ΔFGD\Delta FGD Model
Variable VIF Variable VIF
FGD Rank 1.833 FGD 1.66
Economic Gender Equality 2015 1.276 Rank Economic Gender Equality 2015 1.15
GDP Rank 1.505 GDP Rank 1.493
Table 7: Variance Inflation Factors of independent variables in the economic gender equality changes model.

Table 8 presents the detailed results of both models of changes. Before fitting, we rescaled the ranked variables to have a value between zero and one to allow a better comparison of their relationships, controlling for autocorrelation by including the unranked value of the variable in the previous year. The results of Table 8 are confirmed by ANOVA tests of the FGD rank in the ΔEco\Delta Eco model F=10.195,p<0.01F=10.195,p<0.01, and the non-significant result for the Eco rank in the ΔFGD\Delta FGD model F=0.003,p>0.9F=0.003,p>0.9.

ΔEco\Delta Eco Model ΔFGD\Delta FGD Model
Term Estimate s.e p-value Term Estimate S.e p-value
Intercept -0.011 0.015 0.437 Intercept 0.019 0.010 0.069
FGD Rank 0.039 0.01 <0.01<0.01 FGD -0.002 0.011 0.848
Eco -0.061 0.021 <0.005<0.005 Rank Eco -0.001 0.015 0.94
GDP Rank 0.052 0.01 <106<10^{-6} GDP Rank 0.007 0.015 0.65
Multiple R2R^{2} 0.1501 Multiple R2R^{2} 0.0009
Table 8: Results of robust regression for the model of changes in economic gender equality and of changes in the FGD.

The residuals of the model of changes in economic gender equality are distributed close to normality, as shown on Fig. 22, with a significant Shapiro-Wilk statistic of 0.970.97 and only some small deviations from normality at the tails. Residuals are uncorrelated with all independent variables and do not show signs of heteroskedasticity.

Figure 22: Analysis of residuals of the economic gender equality changes model. The left panel shows the normal Q-Q plot of residuals of the model, and the right panel their histogram. Some minor deviations from normality can be observed in both.

We tested the robustness of the positive association between FGD and ΔEco\Delta Eco in two new fits including the same controls as for the FGD model (VIF of all factors below 5). The results are shown on Table 9, evidencing that the observed association between FGD and ΔEco\Delta Eco is robust to other socio-economic indicators, and to other possible confounds such as Facebook penetration or mean user age.

Term Estimate s.e p-value Estimate s.e p-value
Intercept 0.007690 0.018282 0.67473 0.006294 0.017968 0.72667
FGD Rank 0.024313 0.010963 <0.05<0.05 0.026402 0.010911 <0.05<0.05
Eco -0.072524 0.024797 <0.01<0.01 -0.075975 0.024367 <0.01<0.01
Ineq Rank -0.016436 0.007985 <0.05<0.05 -0.013190 0.007922 0.09833
Pop Rank 0.017917 0.008364 <0.05<0.05 0.019372 0.008090 <0.05<0.05
Mean Age Rank 0.005872 0.012132 0.62919 0.006279 0.011781 0.59494
FB Penetration Rank 0.015300 0.013533 0.26029 0.015453 0.012387 0.21442
GDP Rank 0.024484 0.016641 0.14359
Internet Penetration Rank 0.024763 0.015389 0.11000
Multiple R2R^{2} 0.2009 0.1966
Table 9: Results of the ΔEco\Delta Eco model including additional controls.

We further tested the possible role of other equality indices in the relationship between FGD and ΔEco\Delta Eco. We added all other three gender equality scores as controls (VIF below 5), and repeated the fit. The result, shown on Table 10 shows that the positive association between FGD and ΔEco\Delta Eco is robust to the possible effect of other kinds of gender inequality.

Term Estimate s.e p-value
Intercept -0.029304 0.021736 0.179917
FGD Rank 0.039560 0.016828 <0.05<0.05
Economic Score -0.043693 0.025742 0.091983
GDP rank 0.045928 0.012187 <0.001<0.001
Education Score rank 0.001291 0.013318 0.922924
Political Score rank 0.014899 0.009621 0.123862
Health Score rank 0.004209 0.009031 0.641953
N 139 Multiple R2R^{2} 0.171
Table 10: Results of the ΔEco\Delta Eco model including controls for other gender equality indices.

We tested whether the association between FGD and ΔEco\Delta Eco could be explained by general cultural differences. We combined our dataset with Hofstede’s cultural dimensions: Power Distance Index (PDI), Individualism (IDV), Uncertainty Avoidance Index (UAI), and Masculinity (MAS) (VIF below 5). This limits the analysis to a set of 66 countries common to both datasets, with results reported on Table 11. The positive association between FGD and ΔEco\Delta Eco is still significant, suggesting that the relationship between both variables goes beyond what Hofstede’s model captures in terms of culture.

Term Estimate s.e p-value
Intercept 0.04743 0.03102 0.1316
FGD Rank 0.03792 0.01764 <0.05<0.05
Eco -0.1243 0.03779 <0.01<0.01
PDI 3.217104*10^{-4} 1.759104*10^{-4} 0.0725
IDV 1.754104*10^{-4} -1.708104*10^{-4} 0.3088
MAS -2.490104*10^{-4} 1.546104*10^{-4} 0.1126
UAI 3.148105*10^{-5} 1.366104*10^{-4} 0.8185
N 66 Multiple R2R^{2} 0.362
Table 11: Results of the ΔEco\Delta Eco model including controls for cultural dimensions.

Following the same methodology as for previous models, we stratified our analysis with FGD measured only in a variety of age groups. The results are shown on Fig. 23, revealing the positive and significant role of FGD in the model for all age segments.

Figure 23: Replication of the model stratifying for different age groups.

Our data offers the opportunity to measure the role of FGD in the changes of other gender equality measures, but the indices for Political and Health gender equality have negligible changes between 2015 and 2016. For that reason we can only evaluate the role of FGD for changes in Education gender equality. We find no significant effect of FGD, as reported on Table 12.

Term Estimate s.e p-value
Intercept 0.0613210 0.0613210 <107<10^{-7}
FGD rank -0.0004546 0.0023720 0.848
Edu -0.0606702 0.0099782 <106<10^{-6}
GDP rank -0.0012934 0.0022238 0.562 <0.01<0.01
N 139 Multiple R2R^{2} 0.08
Table 12: Results of the ΔEdu\Delta Edu model.

References

  • [1] Berners-Lee T (2010) Long live the web. Scientific American 303(6):80–85.
  • [2] Hargittai E, Hsieh YP (2013) Digital inequality, ed. Dutton W. (Oxford University Press), pp. 129–150.
  • [3] Brown R, Barram D, Irving L (1995) Falling through the net: A Survey of the ”have nots” in rural and urban America. (United States Department of Commerce).
  • [4] Compaine BM (2001) The digital divide: Facing a crisis or creating a myth? (Mit Press).
  • [5] Norris P (2001) Digital divide: Civic engagement, information poverty, and the Internet worldwide. (Cambridge University Press).
  • [6] United Nations General Assembly (27 March 2006) 60 resolution 252.
  • [7] Group WB (2016) World Development Report 2016: Digital Dividends. (Washington, D.C.: World Bank).
  • [8] Friederici N, Ojanperä S, Graham M (2017) The impact of connectivity in africa: Grand visions and the mirage of inclusive digital development. Electronic Journal of Information Systems in Developing Countries 79(2):1–20.
  • [9] Burke M, Kraut R (2013) Using facebook after losing a job: Differential benefits of strong and weak ties in Proceedings of the 2013 conference on Computer Supported Cooperative Work, CSCW ’13. pp. 1419–1430.
  • [10] Gee LK, Jones JJ, Fariss CJ, Burke M, Fowler JH (2017) The paradox of weak ties in 55 countries. Journal of Economic Behavior & Organization 133:362–372.
  • [11] Jensen R (2007) The digital provide: Information (technology), market performance, and welfare in the south indian fisheries sector. The quarterly journal of economics 122(3):879–924.
  • [12] World Wide Web Foundation (2015) The web and rising global inequality.
  • [13] Wagner C, Graells-Garrido E, Garcia D, Menczer F (2016) Women through the glass ceiling: gender asymmetries in wikipedia. EPJ Data Science 5(1):5.
  • [14] Rizoiu MA, Xie L, Caetano T, Cebrian M (2016) Evolution of privacy loss in wikipedia in Proceedings of the Ninth ACM International Conference on Web Search and Data Mining, WSDM ’16. pp. 215–224.
  • [15] Garcia D, Weber I, Garimella VRK (2014) Gender asymmetries in reality and fiction: The bechdel test of social media in Eighth International AAAI Conference on Weblogs and Social Media, ICWSM ’14.
  • [16] Nilizadeh S, et al. (2016) Twitter’s glass ceiling: The effect of perceived gender on online visibility in Tenth International AAAI Conference on Web and Social Media, ICWSM ’16’. pp. 289–298.
  • [17] Magno G, Weber I (2014) International gender differences and gaps in online social networks in International Conference on Social Informatics. pp. 121–138.
  • [18] Haranko K, Zagheni E, Garimella K, Weber I (2018) Professional gender gaps across us cities. arXiv preprint 1801.09429.
  • [19] Billari FC, Zagheni E (2017) Big data and population processes: A revolution? in Statistics and Data Science: newchallenges, new generations. Proceedings of the Conference of the Italian Statistical Society. (SIS), pp. 167–178.
  • [20] Fatehkia M, Kashyap R, Weber I (2018) Using facebook ad data to track the global digital gender gap. World Development 107:189 – 209.
  • [21] Ojala J, Zagheni E, Billari FC, Weber I (2017) Fertility and its meaning: Evidence from search behavior in Eleventh International AAAI Conference on Web and Social Media, ICWSM ’17.
  • [22] González Cabañas J, Cuevas Á, Cuevas R (2017) Fdvt: Data valuation tool for facebook users in Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems. (ACM), pp. 3799–3809.
  • [23] Pötzschke S, Braun M (2017) Migrant sampling using facebook advertisements: A case study of polish migrants in four european countries. Social Science Computer Review 35(5):633–653.
  • [24] Zagheni E, Weber I, Gummadi K (2017) Leveraging facebook’s advertising platform to monitor stocks of migrants. Population and Development Review 43(4):721–734.
  • [25] Dubois A, Zagheni E, Garimella K, Weber I (2018) Studying migrant assimilation through facebook interests. arXiv preprint 1801.09430.
  • [26] World Economic Forum (2016) World economic forum gender gap report. http://reports.weforum.org/global-gender-gap-report-2016.
  • [27] Katz ML, Shapiro C (1985) Network externalities, competition, and compatibility. The American economic review 75(3):424–440.
  • [28] Hendler J, Golbeck J (2008) Metcalfe’s law, web 2.0, and the semantic web. Web Semantics: Science, Services and Agents on the World Wide Web 6(1):14–20.
  • [29] Hofstede G (2003) Culture’s consequences: Comparing values, behaviors, institutions and organizations across nations. (Sage publications).
  • [30] Ingram D (2018) Factbox: Who is cambridge analytica and what did it do? retrieved April 23, 2018.
  • [31] Cabañas JG, Cuevas A, Cuevas R (2018) Facebook use of sensitive data for advertising in europe. arXiv preprint 1802.05030.
  • [32] Garcia D (2017) Leaking privacy and shadow profiles in online social networks. Science Advances 3(8).
  • [33] Garcia D, Goel M, Agrawal AK, Kumaraguru P (2018) Collective aspects of privacy in the twitter social network. EPJ Data Science 7(1):3.
  • [34] Spaiser V, Ranganathan S, Mann RP, Sumpter DJ (2014) The dynamics of democracy, development and cultural values. PloS one 9(6):e97856.
  • [35] Uteng T (2012) Gender and mobility in the developing world. World Development Report pp. 7778105–1299699968583.
  • [36] Facebook (2018) Facebook ads api. https://developers.facebook.com/docs/graph-api.
  • [37] Metcalf J, Crawford K (2016) Where are human subjects in big data research? the emerging ethics divide. Big Data & Society 3(1):2053951716650211.
  • [38] Ess C, et al. (2002) Ethical decision-making and internet research: Recommendations from the aoir ethics working committee. Readings in virtual research ethics: Issues and controversies pp. 27–44.
  • [39] Araújo M, Mejova Y, Weber I, Benevenuto F (2017) Using facebook ads audiences for global lifestyle disease surveillance: Promises and limitations in Proceedings of the 2017 ACM on Web Science Conference. (ACM), pp. 253–257.
  • [40] World Bank (2016) World bank human developent index. http://hdr.undp.org/en/content/human-development-index-hdi.
  • [41] Chatterjee S, Hadi AS (2015) Regression analysis by example. (John Wiley & Sons).
  • [42] Plummer M, , et al. (2003) Jags: A program for analysis of bayesian graphical models using gibbs sampling in Proceedings of the 3rd international workshop on distributed statistical computing. (Vienna), Vol. 124, p. 125.
  • [43] Koller M, Stahel WA (2011) Sharpening wald-type inference in robust regression for small samples. Computational Statistics & Data Analysis 55(8):2504–2515.
  • [44] Cromwell JB, Labys WC, Terraza M (1994) Univariate tests for time series models. (Sage) Vol. 99.