Showing posts with label HSR 2014. Show all posts
Showing posts with label HSR 2014. Show all posts

Monday, November 10, 2014

Who is the Kevin Bacon of Health Systems Research?

Earlier, I wrote about network mapping at the Third Global Symposium on Health Systems Research. We mapped participants' social ties, collaboration ties, and information seeking ties as reported by them on an online survey. 

Today I will show you the collaboration network generated through this process, as well as the 'real' network of co-authorship relationships generated through data mining. 

We asked: "With whom have you collaborated on health systems or policy research in the past two years." Here's what we got: 

In this graph, 91% of all nodes belong to a single, connected component. And, any two people in this network are connected by four or fewer people. We knew that the HSR community was… well-connected… (insular also comes to mind), but really? 


In network science this is referred to as the 'small-world phenomenon.' Ever heard of 'six degrees of separation'? This refers to the fact that any two people in the United States are connected by six or fewer acquaintances (as confirmed by the famous postcard experiment of Stanley Milgram). Or that everyone in the movie world is connected to Kevin Bacon* by four or fewer connections (you might have also heard six). 


In other words, Kevin Bacon is connected to the most other actors in the fewest steps. Why does this matter? Imagine that you had an important message to deliver to everyone in your network, and that the message travels from person to person starting with one person. Who would you give it to for fastest diffusion? Test your intuition here by trying to slow the spread of an infectious disease in a network. 

Bottom line, you want to find the person who can access different parts of the network in the fewest steps. They sit on paths between the most number of people. They have the highest betweenness centrality

But, our network data were not complete. Only about 70 people completed the survey, so we have plenty of missing edges and attributes for those who didn't complete the survey. 

This begs two questions: 


1. Can we mine citation data to get a fuller picture?
2. Who is the Kevin Bacon of HSR?

To answer the first, yes, we can use publicly available citation data. I searched ISI/Web of Science for all the names of Health Systems Global members (which was our proxy for who attended the conference), and downloaded the names of everyone they have published with since 2012, mirroring the survey question. Using Sci2, I constructed a co-authorship network with our respondents and their co-authors as nodes, and co-authorship relationships as edges/ties. 

This is the network we get: 


Complete co-authorship network of survey respondents


This network has 5576 nodes, representing 1433 publications since 2012. 75% of these authors are connected in one large component, with a relatively dense core but sparser periphery structure. We also see many more components -- a total of 191 unconnected smaller networks. No longer is everyone as connected. Now, like in the real world, everyone is connected by 6 degrees of separation, compared with 4 above. 

Who has the shortest paths to everyone else in the network? Who is our Kevin Bacon?



Co-authorship network of survey respondents, nodes sized by (scaled) betweenness centrality, 
six highest betweenness scored labeled 

John Lavis has the highest betweenness centrality. John** was on the greatest number of paths between other actors. He is in a position to best connect otherwise unconnected parts of the network, and to reach all other nodes most efficiently, all else being equal.

When it comes to ideas, John Lavis is a broker (which is appropriate, as he studies knowledge brokers). He can:


  • Access new ideas from diverse parts of the network
  • Disseminate new ideas to diverse parts of the network
  • Act as a translator between groups



What makes John different than the rest of us? Since 2012, he published 30 articles with 78 different people. That's an average of 2.6 new co-authors per paper! Imagine how many people they are connected to. And the fact that you might not recognize all the names is a good thing for innovation. 

Does he intentionally try to broker the world of HSR? Maybe (we should ask him), but he also has many factors in his favor: he's an expert, so diverse authors seek him out and he's based in Canada where institutions are smaller and thus one cannot rely on within-institutional collaborations alone. 



Did you happen to notice that of the six highest scores in the HSR network above, four are Canadian by birth? Natural peacekeepers?

Back to the big picture. What does this mean for the field of Health Systems Research?


Connectivity is a good thing -- it's how we share and interpret ideas and knowledge, access resources and opportunities, and function as a society -- but there are two concerns with connectivity.

1. More connectivity isn't always better



2. The distribution of connections matter


First, there is a danger of networks becoming overly connected. Multiple, redundant connections with the same groups of people (represented by high network density or many triangles in a network) leads to 'group think' and may stymie innovation. Densely connected network cores begin to operate like an echo chamber.

Second, we need to think of how the connections are distributed across network nodes. At the node level, centrality (i.e., number of connections) equals social capital. We learned that betweenness centrality is one particularly powerful measure of capital, brokerage, and potential influence. The HSR co-authorship network had extremely large variation in people's betweenness scores, meaning that there is a lot of "power" in this network, and it is not evenly distributed.

So high network density is bad? 


No. We need to separate connectedness into two components: density (i.e., the proportion of actor pairs who are connected) and centralization (i.e., the distribution of connections across the network). Density in and of itself isn't bad; sometimes all those repeated connections are necessary. It is the centralization of those connections around one or a few nodes that is not ideal for innovation. We used to think that density and decentralization were at odds but we now know they can co-exist.

Reza Yousefi-Nooraie found that academic medical research teams were more productive if their networks were dense, but also decentralized and open to external networks. A recent study in Nature, corroborating other studies of knowledge-intensive organizations, found that dense networks aid knowledge transfer when the knowledge is complex. The same study also found that team performance was positively associated with the strength of expressive ties (i.e., friendships).

However, I found in Burkina Faso that while dense networks aided the dissemination and spread of research evidence (i.e., complex information), the same density protected the policy status quo and prevented the new ideas from being adopted.

So what does success look like for a HSR network? An ideal academic co-authorship network would be productive (i.e., publish a large quantity of research) but it should also be innovative (i.e., willing to adopt new ideas) and relevant (i.e., research priorities are identified systematically and democratically and research findings are used). What would that look like as a network?


  • Dense, with strong ties, including layered ties based on friendship and trust
  • Decentralized, with brokers at the margins who can access external networks and new ideas
  • Diverse, with varied representation of people who can translate across communities

How do we move towards such a network?

 

As individuals, we can continue to develop strong relationships with co-authors that lead to intellectual productivity. The HSR network, particularly the core, could be more decentralized. We could all be better at developing new collaborations with people outside of our existing networks. Finally, but related to the previous point, we can do better at building diverse research teams, as measured by geographic location, sex, age, discipline, and job role. At the network level, there are interventions that could be applied, but I can't think of examples where co-authorship networks are actively governed. What about research collaboration networks? Is this the role of Health Systems Global (or donors)?


* Note that newer data show that Sean Connery is on more shorter paths than Kevin Bacon.
** John was my PhD supervisor, but this in no way informed the network analysis.

Monday, October 27, 2014

Mapping Networks from HSR 2014


At the recent Third Global Symposium for Health Systems Research, Jeff Knezovich and I asked participants to complete an online network survey. Our aim was to map the networks (social, collaboration, information seeking) of conference participants. We had some technical glitches with the online tool, slow internet access, and the apathy towards completing a survey that is commonly observed. The experience confirmed:
  1. Network mapping, especially in internet-limited settings, should be done off-line;
  2. Network mapping (probably anywhere) should be done face-to-face. Otherwise respondents are unlikely to respond;
  3. One should always pilot their data collection tools!
The idea of mapping the network seemed to resonate, but in total we had only 71 responses (give or take – cleaning the data outputted from the app was more grueling than a hike up Table Mountain!). Nevertheless, let’s see what these data look like. I will be running the analysis in Rstudio, using the statnet suite of packages. Feel free to download the .csv files and work along with me.

Step 1: Make sure the packages are installed

Install the latest version of ergm and sna: install.packages('ergm'); install.packages('sna')
library(ergm)
library(sna)
I also changed the color palette because I don’t like missing node attributes to be colored as black. It just doesn’t look nice. As a note, these colors are also good for color-blindness, and for printing in grey-scale.
col.list <-c("white", "darkblue", "cornflowerblue", "darkorange1", "darkred")
palette(col.list)

Step 2: Import data and convert to network

ONASurveys.com exported the data as an edge-list, and I had to do some extensive work to delete duplicate names, ensure IDs matched, etc. But you can use the cleaned files. There is one for each network, as well as an attributes file that is used for all the networks.
Save the files somewhere and set that folder as your working directory in R.
setwd("~/Dropbox/Dissertation_jan4/Conferences/Cape Town 2014")
attr <- read.csv(file="attr.csv", header=T, stringsAsFactors=FALSE)
social <-read.csv(file="social.csv", header=TRUE, stringsAsFactors=TRUE)
spre <- network(social, matrix.type="edgelist", directed=F)
smat <- as.matrix(spre)
snet <- network(smat, matrix.type="adjacency", directed=F, vertex.attr=attr)
Note that I had to coerce the edgelist into a network object, then into a matrix (to match up with the attributes), then back into a new adjacency-type network object with attributes attached.
summary(snet)
The summary() command shows us the network vertices, edges and density. Note that this is the entire network of all 1515 participants, 90% who did not complete the survey. Let’s plot that network to see what it looks like (and then we’ll get rid of the isolates).

Step 3: Plot the network

plot.network(snet, edge.col="darkgrey", vertex.border="black")

You can see a cluster of activity in the center, with edges in grey. But otherwise this is not a helpful graph. Let’s delete isolates and check out the summary stats.
sno_iso <- delete.vertices(snet, which(degree(snet)<1))
Run the summary(sno_iso) command to see ALL the details, or simply:
centralization(sno_iso, degree, mode="graph")
## [1] 0.1035052
network.density(sno_iso)
## [1] 0.007237999
Still a pretty sparse network! (Although these are incomplete data, so we can’t say much about the actual network density or centralization). The exact question was: “who did you, or do you plan to have lunch or dinner with during the conference.” One can imagine that there are likely clusters of friends/colleagues who are likely to socialize, and fewer connections between these clusters. And what drives the propensity to form social ties? Let’s look at a few graphs before we test hypotheses in ergm models.
Do people socialize with others from their region?
s2coord <- plot.network(sno_iso, edge.col="darkgrey", vertex.border="black")
plot.network(sno_iso,  coord=s2coord, vertex.col="region", edge.col="darkgrey", vertex.border="black")
legend("bottomleft", legend=c("Africa", "Americas", "Europe", "South-East Asia", "Unknown"), pch=21,
       cex=1, pt.bg=c("darkblue", "cornflowerblue", "darkorange1", "darkred", "white"))

Again, it is somewhat difficult to tell with the missing attribute data, but it doesn’t seem as though there is clustering by region. (On a side note, if these data were very important, I could look up all the alters’ regions. For smaller networks this would certainly be worth it).
Is sociality based on similar organization?

Hmm… first of all, we see that most of our respondents are from research organizations. Second, they seem to be more central in the network. Are they more likely than chance to eat lunch with other researchers? We will find out soon. But first, let’s examine by age.

Finally, which nodes are in the most strategic position to broker other nodes? This is measured by betweenness centrality, and can be applied to understand who the brokers are, and how to most efficiently disseminate ideas or information. We will calculate the betweenness centrality scores for all nodes, and then size our graphed nodes according to their betweenness.
s2between<-betweenness(sno_iso, g=1, gmode="graph", cmode="undirected")
plot.network(sno_iso, coord=s2coord, vertex.col="region", vertex.cex=s2between/1500, edge.col="darkgrey", vertex.border="black")
legend("bottomleft", legend=c("Africa", "Americas", "Europe", "South-East Asia", "Unknown"), pch=21,
       cex=1, pt.bg=c("darkblue", "cornflowerblue", "darkorange1", "darkred", "white"))

The most strategically located brokers are from South-East Asia.

Step 4: Construct ergm models to test hypotheses about why conference participants socialize with each other

Ok, now let’s examine these in ergm models. Exponential random graph models (ergm) are a class of logistic regression model that allow us to test hypotheses related to dyads, i.e., network ties/edges. See the statnet website for a list of ergm resources. I highly recommend “Birds of a Feather or Friend of a Friend” by Goodreau, Kitts and Morris (2009) for both a master class in ergm modeling as well as a wonderful application to adolescent friendship networks.
Unlike traditional statistical models, where the covarariates are some function of the units of analysis, ergm models allow us to alo reprensent covariates that are functions of the network itself. I usually build my ergms in two waves: 1. A set of attribute-only models where covariates are tested separately and then added to the final model stepwise if they improve model fit; 2. A set of structural-only models (following the same process as above).
Starting with the attributes, let’s test each in a model with an edges term, which is like an intercept in a traditional regression model.
smodel.02 <- ergm(sno_iso ~ edges+nodematch("region"))
summary(smodel.02)
## 
## ==========================
## Summary of model fit
## ==========================
## 
## Formula:   sno_iso ~ edges + nodematch("region")
## 
## Iterations:  20 
## 
## Monte Carlo MLE Results:
##                  Estimate Std. Error MCMC % p-value    
## edges            -3.66743    0.05216     NA  <1e-04 ***
## nodematch.region -3.02832    0.14465     NA  <1e-04 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
##      Null Deviance: 82741  on 59685  degrees of freedom
##  Residual Deviance:  4375  on 59683  degrees of freedom
##  
## AIC: 4379    BIC: 4397    (Smaller is better.)
The nodematch.region coefficient reports the change in log odds of a tie existing between any two given nodes if they share the same region (as compared to not sharing a region). The exponent of -3.02832 is 0.0484. Two actors are less likely to socialize if they are from the same region. BUT! Big caveat: so much of the alter data is missing, that the majority of ties are between known regions and unknown regions (i.e., different regions).
Let’s look at some models with structural covariates. What are these magical structural covariates? They are underlying social processes which have been documented empirically to occur more than chance alone. Today we will examine transitivity, or triangle formation, which describes the propensity for people to form relationships with ‘friends of friends.’ Transitivity has many implications. Think of triangles, literally, as cliques. Cliques might be fun for lunch, but they are not conducive to exposure to new ideas, innovation, behavior or policy change, etc. In the first model I test whether social ties are more likely to exist if they close a triangle. We expect that a person is more likely to socialize with their friends’ friends.
*A note about missing edge data: While our structural models will not be affected by missing attribute data, they will be affected by missing edge data. We only know the edges of respondents, not the edges of the alters
smodel.06 <- ergm(sno_iso ~ edges+gwesp)
summary(smodel.06)
## 
## ==========================
## Summary of model fit
## ==========================
## 
## Formula:   sno_iso ~ edges + gwesp
## 
## Iterations:  20 
## 
## Monte Carlo MLE Results:
##             Estimate Std. Error MCMC % p-value    
## edges       -5.19997    0.05653      0  <1e-04 ***
## gwesp        0.99650    0.17815      0  <1e-04 ***
## gwesp.alpha  1.18096    0.10367      0  <1e-04 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
##      Null Deviance: 82741  on 59685  degrees of freedom
##  Residual Deviance:  5011  on 59682  degrees of freedom
##  
## AIC: 5017    BIC: 5044    (Smaller is better.)
The gwesp term measures the change in log odds of a tie forming between two nodes given that this tie will close a triangle. Yes, even with missing edges, ties are more likely to exist if they close a triangle between three nodes. This is the result we expected for sociality, but let’s check to see whether this happens with collaboration ties. We would expect that people are more likely to collaborate with their collaborators’ collaborators.
cmodel.06 <- ergm(cno_iso ~ edges+gwesp)
summary(cmodel.06)
## 
## ==========================
## Summary of model fit
## ==========================
## 
## Formula:   cno_iso ~ edges + gwesp
## 
## Iterations:  20 
## 
## Monte Carlo MLE Results:
##             Estimate Std. Error MCMC %  p-value    
## edges       -4.96450    0.06448      0  < 1e-04 ***
## gwesp        0.93190    0.25254      0 0.000225 ***
## gwesp.alpha  1.27732    0.12876      0  < 1e-04 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
##      Null Deviance: 50343  on 36315  degrees of freedom
##  Residual Deviance:  3729  on 36312  degrees of freedom
##  
## AIC: 3735    BIC: 3761    (Smaller is better.)
Indeed, people are much more likley to collaborate with their collaborators (based on our incomplete data).
There are a few ways to deal with missing edge data:
  • We could have asked the respondents to report their alters’ edges (this is typical in ego-network sampling, but its accuracy depends on the relationship being measured)
  • We can remove the nodes we didn’t interview (and thus their edges). This will leave us with a network of complete edges, but not a complete network. I.e., the network we will be left with is not a real network. But neither is the missing edges network…
  • We could try to impute edges based on attribute data. Wait! This is what we’d do if we didn’t know that edges are predicted not just on attributes, but also on network structure! Network dependencies make it difficult to impute edges. Man, these networks! Next time
Finally, we could mine existing data (i.e., citation data, Twitter data) to construct relevant networks. Maybe next week…