I was asked about this last week by a colleague, and now it's hit the blogosphere, so I thought I would publicly leap into a dispute about sexism in science. And make a plea for people to actually look at their data.
This was all started by a group of biologists who have been working at NCEAS on publication biases in ecology (the biggest bias is, of course, that not enough of my papers get accepted straight away). They managed to get their latest results published in TREE.
The received wisdom is that there is a bias against women in science. One area where this might be seen is in acceptance of papers for publication – referees and editors might have a bias (conscious or subconscious) against women. If this is true, the proportion of papers published by women should be higher in journals where the gender of the author is not known.
For this and other reasons there have been suggestions floating around that journals shift to a system of double-blind reviews. At the moment most journals have single blinding: the authors' identities are known to the referees, but the referees' identities are not revealed to the authors (unless the referees wish to do so). In double blinding, the referee doesn't know the identity of the author. Hence, any bias due to gender of the author should be removed. So, if a journal shifts from single blinding to double blinding, the proportion of papers by female authors should increase.
In 2001 the journal Behavioural Ecology moved to double blinding. But did this change the proportion of female authors? Or, more exactly, was there a bias against women that was removed? After all, the proportion of female authors might be changing in the rest of science – the null expectation is that the change in Behavioural Ecology should be the same as in similar journal, rather than there should be no change. So, the group gathered data on the number of papers by male and female first authors from before and after Behavioural Ecology switched to double blinding for five similar journals too. And then they compared the change in proportion of female authors in Behavioural Ecology to that in the other journals.
Err, no.
What they did was to compare the change in the proportion of female authors in each journal to zero. They found was that Behavioural Ecology and Biological Conservation. had increases that were significantly different, but not the other journals. They therefore concluded that there was an effect of double blinding, and that the increase in Biological Conservation must have been due to other factors. Oddly, though, at no point did they seem to make a direct comparison. It is not clear that they looked at the data either. Had they done so, they would have seen this:
The lines show the change from before Behavioural Ecology went double blind to afterwards. The vertical lines are the standard errors. Behavioural Ecology is the thick black line. We see that the proportion of female authors increases in all of the journals, but also that it is greatest in Behavioural Ecology. But is that increase significantly (in any sense) greater than in the other journals? Well, comparing it to zero obviously inflates the estimate of significance, because the other journals are all also increasing.
We can get an idea about if the data show anything with a more focussed analysis. This is also simplified, but I an ignoring some variation, and a more sophisticated analysis (=too much hassle to explain) comes to the same conclusion (and yes, for those who have read the paper, so does including the "don't knows").
What we can do is calculate the difference between the before and after proportions of female authors for the “control group”, and estimate the distribution of differences that would be expected if there was no double blinding implemented. Then we can ask if the difference in the proportion for Behavioural Ecology falls so far outside this distribution that it would be unlikely to explain the change.
These are the differences:Journal Percentage Before Percentage After Difference (%) Behavioural
Ecology23.7 31.6 7.9 Behavioral Ecology & Sociobiology 25.1 26.3 1.3 Animal Behaviour 27.4 31.6 4.2 Biological Conservation 13.8 20.6 6.8 Journal of Biogeography 14.4 16.5 2.0 Landscape Ecology 19.5 23.4 3.9
For the journals in black, the mean difference is 3.65%, with a standard deviation of 2.15%. If these were exact, then there would be a 95% chance that the change for another, similar, journal would be between -0.6% and 7.9%. So, Behavioural Ecology is right on the edge.
But it assumes that the variance is known. In reality it is estimated, and only estimated from 5 data points (i.e. not a lot). If we take this into account, we find that the prediction for a journal would fall between -2.3% and 9.6% (with 95% probability). Now Behavioural Ecology is reasonably well inside the limits. Even someone wanting to do a one-sided test will find it inside.
So, the analysis shows little evidence for any effect of double blinding. But there are a couple of caveats, which could have opposite effects. The first is simply that there is not a lot of data – only 6 data points. We would really need more journals to be able to come to any conclusion. In particular, there may have been some other changes at Behavioural Ecology that could have had an effect.
The second caveat is more subtle. Suppose you were a journal editor, and you introduce a rule that authors have to admit that statisticians are the highest form of life in their acknowledgements. After a couple of years, you notice that the proportion of authors called Fisher has increased. You wonder if this is because of the new rule. So, you compare it with other journals, and find no increase. You therefore declare that Coxes appreciate statisticians, but other people don't. But what about all those other effects you didn't see? What about the changes in numbers of Boxes, Coxes, and Nelders? Humans are very good at detecting patterns, but very bad at judging whether they are random. And using the same data from which you spotted a pattern to assess whether it is real is naughty – of course you're going to see an effect, because you've already noticed it in the mass of all possible things that could happen. Now, I don't know if the authors are guilty here – they don't say how they came to decide to examine this particular aspect of the data, but the introduction is a bit arm-wavy about the effect of double-blinding on sex ratio.
Of course, the solution to both caveats is simple – get more data. Anyone fancy trawling through the literature this weekend?
EDIT: Oops. I should have hat-tipped Grrlscientist for her post, which encouraged me to write this. Hedwig - I'm sorry. Please don't set Orpheus onto me...
Reference
BUDDEN, A., TREGENZA, T., AARSSEN, L., KORICHEVA, J., LEIMU, R., LORTIE, C. (2008). Double-blind review favours increased representation of female authors. Trends in Ecology & Evolution, 23(1), 4-6. DOI: 10.1016/j.tree.2007.07.008
Read more!
Tuesday, 22 January 2008
Gender Differences: Need More Data!
Posted by
Bob O'Hara
at
22:05
9
comments
Labels: papers, statistics
Monday, 15 October 2007
Lunacy and the Planets
Excitation of Lunar Eccentricity by Planetary Resonances -- Cuk 318 (5848): 244 -- Science
I assume this means that lunacy and astrology are linked. Or perhaps that the star signs of werewolves are important.
Powered by ScribeFire.
Posted by
Bob O'Hara
at
16:22
2
comments
Friday, 12 October 2007
It almost makes me feel homesick
In today's ScienceExpress (the online first service of Science):
Widespread Morning Drizzle on Titan
Máté Ádámkovics, Michael H. Wong, Conor Laver, Imke de Pater
Evidently the first man (or woman!) onto Titan should be English. The bad news, though, is that the rain is of methane not water. I haven't tried, but I'm guessing the tea won't taste as nice.
Read more!
Posted by
Bob O'Hara
at
10:29
0
comments
Saturday, 22 September 2007
Re-Cell Cycling Front-loading Pt. II
Sudden emergence holds that various forms of life began with their distinctive feature already intact, fish with fins and scales, birds with feathers and wings, animals with fur and mammary glands.I bring this up because I want to continue where I left off fisking this paper
Sherman, M. (2007). Universal genome in the origin of metazoa. Cell Cycle 6: 1873-1877. Link
Previously, I went through the evidence Sherman put forward to suggest that there was a problem for evolutionary biology, and humbly suggested that it may not be as much as a problem as he thought, and that if he was to make sweeping statements, he might like to support them. So, now let's get on to what he suggests explains the problems he sees. In a nutshell, it is sudden emergence with huge brassy shiny knobs on.
This is what he writes:
Here I propose a hypothesis that answers the questions posed above, and offers experimentally testable predictions. This hypothesis postulates that (1) shortly (in geological terms) before Cambrian period a Universal Genome that encodes all major developmental programs essential for every phylum of Metazoa emerged in a unicellular or a primitive multicellular organism; (2) The Metazoan phyla, all having similar genomes, are nonetheless so distinct because they utilize specific combinations of developmental programs. In other words, in spite of a high similarity of the genomes in phyla X and Y, an organism belonging to phylum X expresses a specifc set of active developmental programs, while an organism belonging to a different phylum Y has a distinct set of "working" programs specifc for phyla Y.(as in my previous post, all grammatical mistakes are in the original. I don't want to indicate them all, so let's just give Sherman a sic note)
In other words, every species had every gene, but not all of them are used. Now, to experienced watchers of ID on the blogosphere, this is a familiar notion, John A. Davison's Prescribed Evolutionary Hypothesis. Davison is a crackpot, but let us not judge Sherman on those grounds. Certainly, Sherman has thought through his ideas more, and is probably a sane, normal person.
So, Sherman's idea is that a Universal Genome appeared in the Cambrian, causing an explosion. Since then, the process of evolution has just been one of changes in the switches in pre-existing developmental programmes. Sherman implies that these changes are not genetic (as all organisms started with the same Universal Genome), so it's not awfully clear how he thinks developmental changes have occurred. Perhaps he should throw some epigenetics into the mix.
So, how did the Universal Genome appear? Ah, Sherman says nothing. It's taken as a given that supporters of ID will say that it can't say anything about the identity of the designer, but this is going one step further - stuff just appeared without mentioning the possibility of a designer. You see, it's Sudden Emergence, only without the mammaries.
Sherman quickly seeks to address what he sees as a fundamental problem: if the universal genome was in all Metazoa initially, why don't we see it in all now? The solution is obvious - genes have been lost over time:
A beautiful illustration of such a loss is a Wnt gene family. In humans, there are nineteen Wnt genes belonging to twelve families. In Hydra. on the other there are two Wnt genes that correspond to two families found in humans. Simple analysis of this finding within the framework of the classical model suggests that additional human genes have developed from ancestral Wnt genes found in Hydra. However, Anemona that belongs to a distinct branch of Cnidaria has eleven Wnt genes belonging to eleven families found in humans. Therefore, it is quite obvious that many Wnt gene families, possibly the entire set, exists in the gene pool of the primilive Metazoan phylum, and various members of Wnt families were lost in different species within this phylum.Waste not, Wnt not.
Accordingly, the proposed model predicts that in various groups of Cnidaria we will find many diverse gene families that function in more advanced phyla.Or that they have also become lost.
Sherman expands on his ideas:
The "Universal Genome" hypothesis does not contradict any well-established data on the genetic evolution (e.g., gene duplications or accumulation of mutations, molecular clock. etc). but suggests that genetic evolution could shape and improve function of developmental programs.It's a pity Sherman missed out all that well know data about large chunks of DNA suddenly appearing in a genome, without any immediate function. I'm a bit sceptical about his suggestion that the hypothesis doesn't contradict data on the accumulation of mutations, so let's see how he defends that:
...Yep, that's how he does it. Once more, make a bold statement, don't give any argument or evidence to support it and move on. For the moment, so shall we.
Sherman fills in a couple of details (including a bit of special pleading for genes have only been lost is some lineages - there must be a special mechanism for their conservation in others. Again, just a a statement, not backed up), and then gets on to the "this is really science" part:
There are two main testable predictions of the presented hypothesis, which are absolutely critical for validation of the model; (1) full or parts of the developmental programs characteristic to higher taxons must be encoded in genomes of lower taxons, and (2) blocks of genetic information encoding these developmental programs in more primitive taxons must be useless in these taxons.What happened to waste not Wnt not? Sherman is saying that primitive organisms must still carry genetic code for more advanced developmental programmes. But he's already accepted that they can lose genes. So, (1) can't be critical. (2) is more interesting - one would need to show that the genetic information was present, but not expressed in the "primitive" taxa. Sherman, of course, doesn't suggest this. Instead, he tells us that the common ancestor of Arthropoda and Chordata didn't have eyes (as we have already seen, this is not quite right), but yet jellyfish have eyes, and they are controlled by similar genes to those in Drosophila. Oh, and as part of this argument, he points out that jellyfish don't have a central processing unit like a brain, so it is unclear how they could process and integrate what they see. Once again, because we are ignorant of something, it can't possibly happen.
Sherman also suggests that we could try and induce development of these more advanced programmes, for example by over-expressing Pax6 in the sea urchin, and seeing whether it develops eyes (and becomes a see urchin, I suppose). It's not clear to me why over-expression would make a difference: we already know that it is expressed in the sea urchin's foot. Without, apparently, an eye being developed. Could it be that other genes involved in eye production are missing? The public needs to know!
We also get another prediction:
Another indication that latent developmental program is present in a lower taxon would be expression of such a program in higher taxons derived from the lower one in a seemingly convergent processes. For example a possible experiment would be to activate development of circulation systems of mammalian or bird types in lizards or even in Xenopus [a frog]. The circulation systems in mammals and birds appear to be very similar, however, they developed from the reptilian system independently in these taxons. Therefore it seems likely that Reptilia possess the program of development of the circulation system of the mammalian/bird type and requires only a minor switch to activate it.This sounds like we would end up with a reptile that was developmentally very confused. Note we're told that Reptilia have the programme for development, but we're not shown any evidence. Where are the homologs and orthologs for the genes? Is Sherman unable to BLAST the Xenopus genome to see if the genes are there? (in case you're wondering, I'm too lazy to chase this up myself).
Sherman suggests his second prediction could be tested by deleting genes and seeing if there is a physiological effect, e.g. the genes for the adaptive immune system in the sea urchin. This makes sense (although you might have to attack the sea urchin with a pathogen afterwards, to see if it responds). Except that it's not clear what one would conclude if nothing happened - Sherman has already discussed the possibility that genes are lost, so he could claim that that's what has happened in this case. In other words, the predictions don't provide potential falsification (although if they were found to be correct, they would provide a powerful verification of the idea).
So, we have a paper that makes some dodgy claims from ignorance that evolution can't explain the Cambrian explosion or the evolution of body plans, followed by and alternative hypothesis which explains nothing that can't be explained by evolutionary biology, relying on gaps in our knowledge to create doubt. And it says nothing about the elephant in the room - how sudden emergence happened. The key part of the hypothesis - how developmental information appeared - is just stated and then left aside.
The whole premise of Intelligent Design as science was that one could investigate design without asking about the designer (because, obviously, that would mean admitting you thought the designer was the god of Abraham). Sherman has taken this to the next stage - he doesn't even mention the possibility of a designer.
I mentioned in my first post that I was suspicious about this paper. That was because the grammar looked odd (the mistakes were in simple grammar, but complex structures were correct). That was before I read this, from the Disco Institute (p20) (pdf):
On the other hand, reading the papers on evolution published in respectedscience journals like Proceedings of the National Academy of Science or Nature, one is surprised at the weakness of the arguments. Indeed, the standards of proof in the field are much lower than in the rest of biology. Such papers would never make it through the peer-review process if they concerned molecular or cellular biology. Of course, there are obvious reasons for such low standards, including the difficulty of testing evolutionary hypotheses through experimentation. But if the theory is based on poor arguments, why have criticisms of it not succeeded in convincing mainstream scientists?Poor arguments? Ha ha ha ha! Now, this guy must be spoofing us, and the DI.
As for the jourmal, it is asking to be spoofed. This is part of the journal's explanation of why one should publish there:
So, the journal takes an hour to decide if a paper is good enough - if it's sent out to review, then it will have to be pretty crap to be rejected. And there is pressure on the referees to comment on a manuscript very quickly. This must have a negative impact on the quality of the refereeing - sometimes you have to go through a paper carefully, and spend time checking references, and also thinking about it - often I need time to work out why I'm not sure about a paper, or to work out what to recommend. The final bulletpoint says to me that the journal is desparate - they want good papers (doesn't every journal?), so they are prepared to cut corners to do so.
Rapid response to presubmission
inquiries (usually within the hour). Most papers are rejected without
external review. (papers send for external peer-review are expected to
be published). During presubmission inquiry, reviewers will be
contacted to accelerate further review. Authors are encouraged to
suggest and decline potential reviewers.
Ultra-rapid peer-review (usually within one-two days)
Papers
rejected from other top journals (e.g., Nature, Science, Cell), if the
authors choose, may be submitted with previous reviews and decision
letters. This allows for the consideration of a paper without sending
for additional review.
The journal is desparate to do things quickly, so it looks like if you wanted to get a dodgy or hoax paper published, this is a good journal to do it with. I hope that is what happened. Otherwise Michael Y. Sherman will have to justify why he can submit a paper in whcih he ignores basic norms of writing science - you know, backing up your argument with evidence. And whether this is a hoax or not, Cell Cycle has to justify how it can publish paper which nobody there has even read properly.
EDIT: Forgot to hat-tip Albatrossity for the pdf. Thanks!
Powered by ScribeFire.
Read more!
Posted by
Bob O'Hara
at
18:18
0
comments
Labels: intelligent design, papers
Monday, 17 September 2007
Re-Cell Cycling Front-loading Pt. I
Last week, my bestest friend DaveScot put up a post at Uncommon Descent (the intelligent design blog of William Demski and other illuminaries) about a paper on front-loading. This is an idea that DS is keen on - that there was an ur-cell that had all of the instructions necessary for all of life, and these were turned on at the right time to produce whatever The Designer wanted to appear. I thought it was worth having a look at the paper, if only to stave away boredom. This is the citation:
Sherman, M. (2007). Universal genome in the origin of metazoa. Cell Cycle 6: 1873-1877. Link
The paper advances a suggestion that goes totally against mainstream evolutionery biology, and is therefore nuts and wrong.
It's almost tempting to stop there, but I doubt anyone would get the joke. So, I'll use a different rhetorical strategy to Sherman, and if I make any grand statements, try to back them up with evidence and argument.
Before laying into it properly, I should state that the paper should never have been published in the form it was. The grammar is awful. I'll comment more on this after I've finished with the text - it makes me a bit suspicious about the whole thing. For the moment, it is enough to point out that the grammatical mistakes in the quotes are in the original, and if I were to acknowledge all of them with the usual sic, this post would look like a vomitorium.
The paper starts by laying out some facts that the author thinks need to be explained:(1) seemingly simultaneous appearence of paleontological remains of all presently existing Metazoan phyla, both simple and advanced; (2) similarities of genomes among Metazoan phyla of diverse complexity; (3) seemingly excessive complexity of genomes of lower taxons; (4) similar genetics switches of functionally similar but non-homologous developmental programs.
Let's take these one by one:
(1) seemingly simultaneous appearence of paleontological remains of all presently existing Metazoan phyla, both simple and advanced. OK, this one is easy - it's boiler-plate creationism. CC300 and CC301 (go to the links for rebuttals).
(2) similarities of genomes among Metazoan phyla of diverse complexity. Grr, now I have to do some work. Sherman points out that some genes (or rather their orthologs) are found in diverse taxa, and not always doing the same thing. He states:...one does not expect to find genes responsible for development of bilateral organisms in primitive Metazoa with radial symmetry. Surprisingly, such genes, e.g., orthologs of hox genes, were found in Cnidaria, and furthermore they are expressed in Cnidaria in an asymmetric manner, as if to define segments in these radial organisms.
Why is this unexpected? Sherman does not explain. Perhaps he hasn't heard of common descent. Or co-option, where genes that have one function are used to do something else (any intelligent intelligent design supporter should know that the bacterial flagellum took some of its structure from the Type III Secretory System). Sherman does discuss genes changing function over evolutionary time:A possible response to these arguments within the classical model would be a suggestion that the genes responsible for eye development in Arthropoda or vertebrates serve different functions in lower taxons (so-called gene sharing). In fact, several examples of gene sharing have been described, e.g., recruiting of small heal shock proteins to serve as crystallines. These examples, however, are exceptionally rare, and it is unclear whether they indeed can be responsible for making de-novo complex developmental programs serving unrelated functions.
So, Sherman, if this is exceptionally rare, why does Conway-Morris declare co-option to be "rampant" (pdf)?...co-option and redeployment are rampant both in a developmental
context (e.g. Eizinger et al., 1999; Heanue et al., 1999 (see also Relaix and Buckingham, 1999); Merlo et al., 2000; Damen, 2002; Locascio et al., 2002; Lowe et al., 2002; Fabrizio et al., 2003) and in related topics such as those concerned with enzymatic pathways (e.g. Peregrin-Alvarez et al., 2003).
Look, look! Conway-Morris acts like a fusty old academic and gives citations! Curse the man for making it so easy to check his assertion!
Of course, it could be that Sherman is unaware of this work, because he hasn't read this paper. Except ... it's cited in his paper. Not that that means much (find the Know-Thine-Own-Self Results).
It's around here that Sherman makes a comment so factually wrong even I spotted it. He writesIn fact, many of the regulatory genes were lost later in evolution, and are not present in Drosophila or C. elegans, e.g., hedgehog gene,5 indicaling that their presence is not necessay for development and life of vey complex Arthropoda.
Um, but hedgehog is found in Drosophila. It must be - it has the requisite silly name. Even funnier, it was discovered in the fruit fly! Oh, and the same page shows that there are genes similar to hedgehog in C. elegans.
Where were we? Oh, next point...
(3) seemingly excessive complexity of genomes of lower taxons; An immediate problem here is how one defines complexity. PZ Myers has a nice essay on this. But let us proceed. Sherman points out that the sea urchin, which apparently is primative (I guess this means it doesn't know how to eat spaghetti properly), has a whole suite of genes involved in eye development: While the presence of the opsins could be explained by their possible function in a simple light sensing, sea urchin has the entire set of orthologs of major genes involved in the eye development ... Therefore, it appears that information on the eye development is encoded in the sea urchin genome, while no eye is actually developed, and thus the genetic information seems to be excessive.
He also points out that the sea urchin has the genes for an adaptive immune system.Yet, sea urchin does not have antibodies, and possibly lacks adaptive immunity in general. Genes that are seemingly useless in sea urchin but are very useful in higher taxons exemplify excessive genetic information in lower taxons.
Or perhaps they exemplify our lack of knowledge about the sea urchin. Now, I know that Sherman is at Boston University Medical School, but I have no idea what he does there. I'm not, though, going to infer that he's useless. Or at least not on this basis.
One can't simply point at a gene and say "we don't know what it does in this organism, so it must be useless". In the paper Sherman cites about the presence of the adaptive immune system genes in the sea urchin, the authors point out that we know very little about the immune systems of most species. To nake his case, Sherman has to show that these genes are not used by the sea urchin, e.g. show that they are not expressed. Put the promotor next to GFP, transform it into the sea urchin, and watch to see when GFP is expressed (it glows green - very pretty). Perhaps the genes are active in the early part of the development of the visual sensory system, and this has been well conserved. Or, again, we could posit co-option of genes from one function to another. Just slap it in, and see when it glows!
There's a recurring theme in this paper - the author makes bold statements and utterly fails to back them up, with evidence, argument or citation. I'm not a developmental geneticist, so I don't want to follow all the claims up, although a few are certainly false (e.g. hedgehog above). Others may be correct, but how are we to know? Can we assume divine revelation?
Right, last point.
(4) similar genetics switches of functionally similar but non-homologous developmental programs. Oh, now the evo-devo people are going to love this:A distinct set of data that call for a novel approach to evolution comes from comparison of genes that control functionally similar genetic programs in Chordata and Arthropoda. There appears to be a high degree of similarity in some of these genes. A classic example of such similarity is Pax6 gene that controls development of visual systems. According to all current accounts, a common ancestor of Chordata and Arthropoda was a very primitive organism that lacked eyes, and therefore the evolution of eyes in these groups was convergent.
All current accounts? Well, I guess Wiki isn't current then. Neither is Sean Carroll, who suggests a common ancestor had proto-eyes at least (p123 of From DNA to Diversity). Or perhaps Sherman is making a bold claim without any evidence. Again. Carroll et al. suggest that the development of the proto-eye that the common ancestor had was under control of Pax6, as it is so conserved. Sherman appears to be unaware of the idea of common descent. I hear it's a rather popular theory that's doing the scientific rounds.
So, to summarise where we've got to: Sherman has suggested that there is evidence for a problem, but has been unsuccessful in actually providing it. He goes on to suggest an alternative, which is a delight to snigger at. But that, dear fools who have gotten this far, is for another day.
Read more!
Posted by
Bob O'Hara
at
18:29
0
comments
Labels: intelligent design, papers
Friday, 17 August 2007
The Wonders of Modern Technology
I think I submitted a paper today, but I'm not sure. At least I tried to.
The problem was the electronic submission. I was sending the manuscript (it's on Bayesian variable selection, for anyone who cares) to a statistics journal, so they like things in LaTeX. I had prepared the manuscript that way (my co-author is a LaTeX guy as well, so it was a way of learning how to use it). But they wanted the initial submission as a .pdf.
LaTeX writes documents out as Postscript files, and the pdf maker on TeXmaker doesn't like my figures. So, I had to make a Postscript version, then found out I couldn't convert to pdf on my Windoze machine, so I moved over to my Linux box, and used ps2pdf to convert it. The manuscript looked fine in Acrobat. Then I tried to submit. OK, first I realised we had forgotten to write an abstract, so had to go back, do that, write a postscript version, move to Linux and convert to a pdf. When I (finally) uploaded the pdf, the journal's software tells me it's not a pdf! Wha? It looks like one to every pdf reader I can lay my hands on (which is 3 or 4 - isn't Linux great?). But for some reason...
In the end I just emailed everything to them, with an apologetic note, asking how I could convert my pdf to a pdf. They haven't answered.
Read more!
Posted by
Bob O'Hara
at
22:13
3
comments
Labels: papers, silliness, submissions
Thursday, 26 July 2007
Reactions to the world's largest glider
A couple of weeks ago, a paper came out online in PNAS, describing the aerodynamics of Argentavis magnificens, a big bird (presumably it had yellow feathers too).
The serious report on this work was done a few weeks ago, but my, err, friend Henry Pihlström has been looking at the paper too closely.
Henry pointed out some details in Figure 4. Here is the figure:
And, for those who don't stare obsessively at journal figures (or at least not as obsessively as Henry evidently does), here is Fig. 4C in detail:
I think the man's reaction is understandable - you wouldn't see that coming towards you on Sesame Street.
Reference
Chatterjee S, Templin RJ, Campbell KE. (2007) The aerodynamics of Argentavis, the world's largest flying bird from the Miocene of Argentina. Proc. Natl. Acad. Sci. 104: 12398-12403.
Read more!
Posted by
Bob O'Hara
at
13:07
0
comments
Labels: ornithology, papers, silliness
Wednesday, 18 July 2007
African or European?
Sometimes I feel obliged to make the obvious joke, just so that nobody else has to embarrass themselves. So it is with this paper.
I'll look at it properly later, but for the moment it suffices to say that they look at the flight speed of different birds, and try to explain that by allometry (shape) and other things. The important point is that they miss of one species: the fully-laden swallow.
Shame on them!
Reference: Alerstam T, Rosén M, Bäckman J, Ericson PGP, Hellgren O (2007). Flight Speeds among Bird Species: Allometric and Phylogenetic Effects PLoS Biology 5(8), e197 doi:10.1371/journal.pbio.0050197
Read more!
Posted by
Bob O'Hara
at
09:48
2
comments
Tuesday, 17 July 2007
On the growth of literary yeast genes
I think one reason I enjoy working as a statistician so much is that we can use the tools of my trade to tease out patterns in data, and actually learn something about the world we live in. One aspect of this is the realisation that if one is trained in statistics, then one is aware of how much more can be extracted from the data. Oh, and also spot subtle mistakes. So, whilst I feel statistics is a service industry: there to help others in their work, there is also a strong element of either pride or egotism in my work. In the light of this...
A couple of computer scientists have just had a paper released in PNAS on "Temporal patterns of genes in scientific publications" (all figures below are from the paper). They looked at how references to genes changed over time in the scientific literature, in particular yeast genes, and whether there was variation in the way different genes were treated. I suspect this was done because it looked like a fun thing to do (a motivation I thoroughly approve of), but it does have implications for the way science is done. We might expect that some "super genes" are intrinsically more important, and hence there is more work on them than on others. Or the patterns of investigation might just be a result of following the herd: people study genes just because others are studying them. Now, a paper like this can't give a full explanation for the patterns of investigations, but it can give some useful information: it can show what the patterns are, and so help focus a more detailed investigation on the interesting aspects of what's going on.
What they did was to get their computers to trawl through the scientific literature and find all the mentions of thousands of yeast genes between 1975 and 2005. Having done that, they then asked how the pattern of references to these genes evolved. The data for two genes look like this:
ACT1 is the most popular gene, whereas PET54 is more typical. The solid lines give the actual data, the dotted line is from a simulation. They assumed a simple model where the number of references to a gene depended on three things:
So, they're imagining a system where work on a gene starts occurs randomly (this is the background rate), but the more work there is on yeast genes, the more likely work on a new gene will start. Makes sense: for whatever reason, someone gets a grant and starts to work on a gene. And this can be seeded by work on another related gene that is part of the same biochemical system. Once work on a gene is started, it should beget more work on that gene; it's good scientific practice, every answer brings up 3 more questions (and hence three more grant applications).
This model wasn't quite good enough though, so they added a saturation term: looking at the data, eventually the rate of citations slows down (for those who are interested, they use the Maynard-Smith and Slatkin model of density dependence). They then fitted this model to the data, assuming the same parameters for all genes (*). At about this point I think "whooo, dodgy". Determined to thwart me in proving my superiority, they then show that this model predicts the distribution of references pretty well. They suggest that what this is showing is that the assumption that the parameters are the same is reasonable. In other words, the genes are behaving equivalently, and there aren't any genes that are more important (and hence more highly cited) than others.
What this implies is that there is nothing intrinsically more important about any one gene as compared to any others: they are behaving as if they are equal, and the reason any one gene is researched more is just happenstance. Once genes become popular, they get researched on more. I guess this could be explained simply because there are more questions that can be asked about a better studied gene.
A few things bother me about this research, though. First, the interpretation of the parameters is off. The authors claim that, if there is no saturation, the growth rate of citations is 0.23, the sum of the growth rates for the gene effect and the general rate of gene reference. Well, um, OK. Let's look at the equation (ignoring saturation, for simplicity):
Ni,t = 0.028 P*t-1 + 0.20 Pi,t-1 + 0.005
Nt is the number of new references about a gene in a year, P*t-1 is the total number of references about all genes (including gene i) up to year t-1, and Pi,t-1 is the total number of references about gene i up to year t-1. So, the growth rate is only 0.23 (=0.028+0.20) if the number of citations for any one gene (Pi,t-1) equals the total number of citations of all genes (P*t-1) (and so the total number of genes is... work it out for yourself). There are over 4000 genes in the data base, so the strength of effect of the total genes should be much higher: on
average, the effect of P*t-1 shoud dominate.
They compound this error by claiming that most of the growth is caused by the gene itself. The model above says that an increase in one reference to one gene in one year leads to an average of 0.2 references in the next year, whereas an increase in one reference to a yeast gene in general leads to 0.028 references (on average) the next year. Look like the effect of a single gene is stronger, doesn't it?
No. There are 4000 genes, so there is more variation in the total number of references of genes. If we only had 100 genes, all equally cited, then if they all increased by 1 reference in a year, then Pi,t-1 would increase by 1, leading to an increase in reference due to this of 0.2, but P*t-1 would increase by 100, so causing an increase of 0.028*100 = 2.8. i.e. over 10 times higher.
Statistically, the point is that the interpretation of regression coefficients depends on the variation in the covariate: one can get larger coefficients simply by changing the measurements from kilometres to nanometres. Substansively, the point is that the dynamics are probably not being driven by single genes, but by the overall rate of growth in yeast genetics. This could be because yeast genetics overall is experiencing growth, or because there has been an inflation in papers generally in science. But it suggests that the dynamics are even more neutral: work on a gene does not beget more work on that same gene. But, I am not convinced that we can conclude this, because...
My other concern is that it's not at all clear to me that the model is working well. One problem is that the model seems to predict too few intermediate genes. This is their plot:
The dots are the observed frequency, the crosses are the simulated frequency. In the middle of the plot, the dots are all above the crosses, so too few genes are being predicted with intermediate frequencies. I'm not sure what this means: usually one would expect the opposite pattern, as variation in the parameters make some genes more extreme. My initial guess is that the rate at which genes appear is not homogeneous: if genes appear at a greater rate in the middle of the time period studied, then this is what we would see. The simulation would spread these genes out over the whole time period, so there would be more genes appearing early (so having more citations) and late (so having very few citations).
Unfortunately, I'm guessing about this, which bring me on to one of my big Statistician Are Important points. The way the model is checked is very simple, so it could be missing a lot of problems. My main worry is that the data were collected at the gene level, so I would like to see how the model predicts the behaviour of individual genes. In the first plot above we see a couple plotted, and it looks like the prediction for ACT1 is poor: it badly under-predicts the behaviour from about 1995. However, this is only one simulation: what about the rest? If they all do this, then there is a problem with predicting the popular genes, i.e. the ones that yeast biologist have decided are important.
There is a lot of data, but something as simple as plotting the residuals (the difference between the observed values and those predicted by the data) against the predicted values is often very telling. For example, if the model fit plot above was plotted as a residual plot, it would give a curve that would be frowning at us, whereas a good residual plot shouldn't have any visible pattern. This could simply be done for all the data (giving one huge plot!), or for subsets (e.g. sub-setted by popularity). I suspect all sorts of problems would be seen: some important, others trivial and ignorable.
So, OK. I'm not a great fan of the analysis. What would I do? I think the first thing would be to look at the data. There is too much to look at it all, but it might be possible to cluster the genes into groups according to their publication record, and look at representatives from each group. From that, I could get a feel for how they are behaving: are most increasing exponentially? Do many reach a plateau? Do they decrease in publication (as researchers get bored and move on to other genes)? Do different genes become popular at different times? Or do all the plots look similar? This would guide me in deciding what sort of model to fit.
Then I would look at fitting a model with different parameters for each gene. I might drop rare genes first: the ones that only get mentioned once or twice. Any estimates for them will be crap anyway. Now, this leads to a lot of parameters, but we can use an approach called hierarchical modeling: basically, assume that the parameters are related, so knowing about 20 parameter values tells us about the 21st. This gives us a direct way of measuring how similarly the genes are behaving: if the parameter estimates are almost the same, then the genes are acting the same. We can even ask where the differences are: is it that genes are appearing in the literature at different times? Or are the growth rates different - i.e. do some genes spawn 10 publications per publication whilst other only spawn 1 or 2?
Once this has been done, we can then check if the model actually works, i.e. if it actually predicts the data well. If it doesn't, then we have to do the process again, based on what we have learned. This can be something of an iterative process, but eventually it should converge to a reasonable model. Hopefully.
The advantage of the approach I'd envisioning is that it becomes more of a conversation with the data (or an interrogation, if it doesn't want to behave). The process is one of learning what the data is telling us, using the plots and models to pick out the important information. We then need to interpret this information (correctly!) to be able to say something about the system - i.e. whether the genes are behaving the same, or what are the differences in rates of reference. And also the extent to which work on one gene is driven by previous work on the same gene, or whether yeast genetics is working by looking at genes in general. One could even try to ask questions about whether this has changed in time: has there been a shift from focus on a single gene to a more general, holistic approach (i.e. do the regression parameters in the equation above change over time. More hierarchical models)? These are all questions that could be tackled: the techniques are there, as long as the data is prepared to cooperate.
Stepping back a bit, and trying not to trip over the cat, this shows a response I have to many papers. The questions are interesting, but the data analysis and interpretation can be improved. For this paper, the problems in the analysis are fairly typical: I suspect the authors are not aware of the techniques that are available. This is understandable, as statistics is its own field, and I would be as lost in, say, yeast genetics. It's because of this that statisticians are needed: even if it's just to have a chat to over coffee, or a pint of Guinness. The interpretation aspect is more of a problem for me: it does not need a great technical knowledge to understand what is wrong, just an appreciation of what the equations mean.
Overall, could do better. But I hope I've shown how to do better. Constructive criticism, and all that.
Reference: Pfeiffer, T. & Hoffmann, R. (2007) Temporal patterns of genes in scientific publications. Proc. Natl. Acad. Sci., 104: 12052-12056. DOI: 10.1073/pnas.0701315104.
* Technical note: they used a general optimisation algorithm (optim in R) to do this. However, given the values of alpha and PS, the model is a generalised linear model with an identity link. So, they could have used used this fact, and just written a function that would optimise alpha and PS, fitting the model as part of the function. Actually, I think I would have started with a density dependence function with one less parameter: even easier!
Read more!
Posted by
Bob O'Hara
at
09:45
1 comments
Labels: genetics, papers, statistics