Showing posts with label academia. Show all posts
Showing posts with label academia. Show all posts

Sunday, 10 February 2008

Fame, Journals, and Blinding

A couple of weeks ago I blogged about a paper on the (possible) effect of double blinding on the bias against female authors. The paper had stirred up a few other comments, which got me thinking a bit more about why I'm a bit sceptical about double-blind reviews. Then I started thinking too hard, and ended up playing around with a simple model.

It's generally agreed that there is a bias towards better-known authors, so that a well-known author is more likely to have a manuscript recommended for acceptance than someone unknown. The argument for double-blinding is that it removes this bias, because the referee doesn't know who the author is. The problem with this is that it is often possible to guess who the author is (hell, it's sometimes possible to guess who a reviewer is) - a study 10 years ago (Cho et al. 1998) found that reviewers could work out the identity of the authors in about 40% of cases.

Presumably the authors who are recognised are the better known ones. We therefore have a situation where fame (whatever it is exactly) affects both whether a paper will be recommended for acceptance, and also whether the authors will be recognised. What effect does this have on the pattern of acceptance? Rather than just indulging in arm-waving, we can build a model, and indulge in arm-waving with numbers!

The model is simple, but hopefully captures the main points. Each author of a manuscript (for simplicity I will assume that each paper only has one author) has a fame. If the author's identity is known to the reviewer, then the probability of acceptance increases with their fame (the solid black line below).

If the reviewing process is double binded, then the probability that the author is recognised increases with the fame (the red dotted line). Note that it starts from a lower point, but increases more rapidly. if the author's name is not recognised, then the probability of acceptance is equal to the minimum probability. This is the the solid red line.


The technical details are below, for those who care. I have also scaled the probabilities of acceptance, so you can see them.

What does this show? Well, if you're a nobody, then the double-blind process means that you do as well as anyone else who isn't recognised, i.e. all but the famous. The famous do well under both systems, as they're recognised anyway. The people who lose out are those in the middle: the ones who are just starting to make a name for themselves, but are yet to be well known. With single blinding, their fame is enough that it helps them. Under double-blinding, though, they are not famous enough that they are recognised, so they are treated the same as a novice.

What this suggests, then, is that double blinding doesn't remove the biases: it just shifts them. So, the very famous actually do better under double blinding, as do the very obscure. Playing around a bit with the model suggests that the general result is robust, but it depends on the probability of recognition starting lower and having a steeper slope.

This is a model, using numbers that were plucked out of the air. But how does it compare to reality? My guess is that the effects are not as severe as shown here, but what is needed is data which can be used to estimate the parameters of the model. In the mean time, I'm not going to submit to any double-blind journals until I have my FRS.



The Maths
Fame, f, is uniformally distributed between -1 and 1. The probability of acceptance for a fame f, pa(f), is modelled like this:


if the identity of the author is known, otherwise it is the minimum value. If the manuscript is double-blind reviewed, then the probability that the reviewer correctly recognises the name of the author, pr(f), is


If a manuscript is reviewed double-blind, the probability that it is accepted is proportional to

pr(f)pa(f) + (1-pr(f))pa(-1)

The final probabilities are normalised, so that they sum to 1, by dividing by the sum of the probabilities.

References
Cho, M.K. et al. (1998) J. Am. Med. Assoc. 280, 243–245.

Read more!

Thursday, 31 January 2008

Big Grant Deadline Today!

Today is the deadline for applications to the Finnish Academy, so scientists all over the country are busy writing abstracts and trying to find an amusing acronym for their project. The bad news for them is that they are up against fierce competition.




Yes, the beast will soon send off the description of his project on Integrated Pest Management For The Removal of the Coarse-haired Wombat (Vombatus ursinus) from Finland.

35% is a reasonable proportion of the grant for me to charge as an overhead, don't you think?
Read more!

Monday, 31 December 2007

A physicist stumbles into a statistical field


By a virtual game of Chinese Whispers (in which no Chinese were involved), I found out about a paper on arΧiv where a poor unsuspecting physicist wanders into a curious part of statistics. I'm actually something of a bystander in this area, but it's not going to stop me commenting on it.

OK, so the paper is by a guy called Bruce Knuteson, from MIT. He's interested in working out the scientific worth of a piece of empirical work, and being a physicist, he wants to measure it.

So, the first problem is to decide what worth is. Knuteson decides to measure it in terms of "surprisal", i.e. how surprised we are by a result. So, if we collected data, and got a result (say, a measurement of a parameter) xi, how shocked would we be by it? From this, Knuteson decides that ...


The scientific merit of a particular result xi should thus (i) be a monotonically decreasing function of the expectation p(xi) that the result would be obtained, and (ii) be appropriately additive.


and so suggests -log(Pr(xi)) as a measurement, as it has these properties. He then suggests that the worth of an experiment can be estimated as the expected value of this, i.e. the sum of -Pr(xi)log(Pr(xi)). This is a measure called entropy: something beloved of physicists and engineers, but rather opaque to the rest of us. The idea is that a larger entropy will mean that the experiment is better - we will expect more surprising results.

But is this a good measure? Perhaps a good way of tackling this is to view it as a problem in decision theory. How can we decide what is the best course of action to take when we are uncertain what the results will be? For example, if we have a choice of experiments we can carry out, how can we decide which one to do? To do this we first need to define "best". This has to be measured, and the numerical value for each outcome is called the utility, U. This might, for example, be the financial gain or loss (e.g. if we are gambling), or might be something more prosaic, like one's standing in the scientific community (however that is measured. h-index?). All the effects of each action, both positive and negative, go into this number. So, for example, we would include the gain in prestige from publishing a good paper, and the cost (e.g. financial, or the effect on our notoriety if the results are a turkey). The second part of the decision analysis is to give a probability for each outcome, so for action A the probability might be 0.3 that we get a Nature paper, and 0.7 that we get a Naturens verden paper. For action B it might be 0.9 and 0.1 respectively. We then calculate the average utility for each action, i.e. sum the probability of each result multiplied by the utility for that result.

This is what Knuteson does to get his measure. The problem is that his only utility is surprisal, and in general this doesn't make sense. Two things are missing. Firstly, there is no cost element. So, if we want to measure the time it takes an apple to fall on a physicist's head, it makes no difference if we pay a couple of students $1 or £30,000,000 to do it. The second problem is that there is no measure of scientific worth. Finding out if the next toss of a €1 coin is treated exactly the same as finding out if the Higgs boson is green.

This leads to clearly nonsensical results. If there are only two possible outcomes of an experiment, then the maximum expected surprisal occurs if the probability of one is 0.5. Therefore the optimal experiment is one with this property. For example, tossing a €1 coin. According to Knutsen, then, we should fund lots of coin tossing experiments (hmm, there's an Academy of Finland application deadline coming up).

The second thing that is missing is where the probabilities come from. These are probabilities of outcomes that are not observed, so in general they cannot be measured (without doing the experiment...). Therefore one has to assign them based on one's subjective opinion. Now we are on familiar Bayesian ground, and is something that has been argued about for years. But here I think Knutsen can use a sneaky trick to sidestep the problems. Put simply, he could argue that in practice the estimation of merit is made by people, so they can assign their own probabilities. If someone else disagrees, fine. This way, it is clearer where the disagreement lies (e.g. which probabilities are being assigned differently).

So, estimating the merit of a piece of work before it is done can be problematic (and I haven't touched on comparing experiments with different numbers of possible outcomes!). But Knutsen develops his ideas even further. How about, he asks, worthing out the merit of an experiment after it has been done?

Before doing this, Knutsen sorts out a little wrinkle. It is not generally the experimental results themselves that are of interest - it is how they impact on our understanding of the natural world. We can think about this in the way that we have several models, M1, M2, ... Mk, for the world (these might correspond to theories, e.g. that the world is flat, or that it is a torus). The worth of an experiment could then be measured in terms of how it changes what we learn about these experiments, i.e. how Mj changes with the data, xi. This can simply be measured as the entropy of the models, the Mk's, rather than the experimental outcomes.

Knutsen goes through the maths of this, and finds that the appropriate measure of the merit of an experiment is a measure of how far the probabilities of each model are shifted by the experiment. To be precise, it is a measure known as the Kullback-Leibler divergence (I will spare you the equations). Now, this again is something that is familiar. A big problem in statistics is deciding which model is, in some sense, best. This can be done by asking about how well it will predict an equivalent data set to the one being analysed. After going through a few hoops, we find that the appropriate tool is the K-L divergence between the fitted model and the "true" model. Of course, we don't know the true model, but there are several teaks and approximations that can be made so that we don't need to - it is the relative divergence of different possible models that is important. The result of this is a whole bunch of criteria that are all TLAs with IC at the end - AIC, BIC, DIC, TIC, and several CICs.

The optimal experiment is the one which will maximise the difference between our prior and posterior probabilities of the different models (yes, Bayesian again). The idea is natural - the greater the difference, the more we have learned, and hence the better the experiment is. Of course, we still have the same problems as above, i.e. assigning the probabilities, and getting the utility right, but we are in the right area. Indeed it turns out (after browsing wiki) that the idea is not original - the method proposed by Knutsen is the same as something called Bayesian D-optimality. And (after reading the literature), the idea goes back to 1956!

So, does this help? For the general problem of estimating scientific merit, I doubt it. There are too many problems with the measure. It may be useful for structuring thinking about the problem, but in that case it it little different from using a decision analytic framework.

In experimental design, it is of more use, but then the idea is not original. The other area it might be useful is in summarising the worth of an experiment for estimating a parameter, such as the speed of light. There will be cases where physical constants have to be measured. Previous measurements can then be used to form the prior (there are standard meta-analysis methods for this), and then the K-L divergence of several experiments can be calculated, to see which gives the largest divergence. This is some way from the ideas Knutsen is thinking about (he explicitly rejects estimating parameters as being of merit!). But I think more grandiose schemes will die because of naysayers like me nagging at the details.

Reference
Lindley, D.V. (1956). On a Measure of the Information Provided by an Experiment. The Annals of Mathematical Statistics, 27, 986-1005.
Read more!

Monday, 8 October 2007

Back again


I've been away for a few days on a course in Estonia, on management and supervision of academic research. This has spawned a few random thoughts:

1. If you take your laptop, remember the power cord. It's more important than the mouse.
2. As I might have expected, some of what we were told I had already worked out, but some I hadn't. It's nice to know which is which. There were also a few ideas that I hadn't thought about.
3. Sharing a room with someone who snores is not a good idea. And going down with 'flu isn't sufficient for revenge.
4. Management theory can actually be useful - obviously there is a lot that is common sense, but having it organised does help. In particular, there was a discussion about project management. In classical project management, one sets out the stages that are needed to complete the task (say, build a bridge), and set out the schedule by working backwards from completion, to decide when each task should be scheduled. I guess this is used all over the place, and is one of the reasons western society works (shocking, isn't it?). Working out how to take an organised approach to this sort of task, and then how to teach this to managers is actually a good thing. IOW, MBAs are not necessarily useless.
5. Thanks to one of the speakers, I'm now reading text with a Welsh/Swedish accent. Should I seek medical help?
6. I hope my students don't mind doing a lot of writing.
7. If you find yourself on the M/S Star between Helsinki and Tallinn, don't try the pizza from the fast food place. It's awful (and remember, I'm English, so I know bad food) - a slab of congealed artificial cheese with chicken and battery-farmed pineapple. It was called Hawaij, presumably because they feared a law suite for defamation from the good people of Hawaii. The people of Four Seasons do not seem to have as much of a reputation for litigation.
8. I don't know what happened in the rugby World Cup on Saturday. I think the media must be lying to us. Ah well, an England-France semifinal, just like last time.

Remind me, who were in the other semifinal in 2003?
Read more!

Saturday, 28 July 2007

The Academic Cult

Does anyone know who Bug Girl is? We have evidence that she needs to be reprogrammtrained.

It's all for her own good. Without retraining, she'll never be able to get another grant.

Read more!