Please, please, please, please... document your data.
It's not fun to get a matrix without any column or row labels. Like my old physics teacher used to say when somebody proudly proclaimed that the answer to an exercise was 5.029, "5.029 what? Elephants?". In physics, a quantity is not meaningful without a unit, and in data management, a matrix is not useful without labels.
It's not fun to know nothing about any preprocessing of the data. Has it been normalised? I guess I can check the mean and standard deviation, but what if it's only close to 0 and 1, but not exactly so? Was that some special normalisation method? Hell if I know.
It's not fun to know nothing about the experimental setup. Maybe you told me every second data point is a wildtype. Does that mean that these were results from two-colour arrays? Or two singe-colour arrays? Are the wildtypes from the same time point as the mutants? Come to think of it, are these even time-course data?
It's not fun to know nothing about what the biologists* want you to find. Are they looking for similarly expressed genes or for regulators? Would they rather have a network of the knocked-out genes or of all genes? Is it worse if I give them false positives or false negatives?
So please, please, please, document your data. Maybe some time in the future I will tell you about fun things called wikis and databases, but for now even a text file would do.
*Or insert other applied science here.
Monday, October 13, 2008
Monday, September 29, 2008
The View from the Foothills
I'm finally there. Five years of hard work, first as an undergraduate and then as a Masters student, have paid off. Last Monday, I was granted my rightful place in academia... on the lowest rung of the ladder.
Yes, PhD students are a dime a dozen, even in my small institute, and despite the excitement of starting research in earnest, I can't help but feel slightly apprehensive. This may just be the result of reading too many PhD comics, but a tinge of anxiety is setting in. What if my advisor turns out to be a workaholic? What if I can't finish in the required 3 1/2 years before my funding runs out? What if my office mates are insane (they're not, I think) or my experiments all fail?
Then I remember that a million students have survived their PhD just fine before me and a million will again. I may not have the prettiest office (in fact, drab is not an inaccurate description), but at least I'm not sharing with 14 other people like my flatmate. My supervisor has only been nice to me, despite the bollocking that he gave his other PhD student last week. And my project, even though it looks daunting from here, will rest on the foundations of my Masters project, meaning that I have a reasonable idea of where to start.
So, the base camp has been established in the foothills of Mt. PhD. Only the future will show if I scale the summit triumphantly, or freeze to death in a crevice somewhere. Boy, that metaphor took a bleak turn, didn't it?
Yes, PhD students are a dime a dozen, even in my small institute, and despite the excitement of starting research in earnest, I can't help but feel slightly apprehensive. This may just be the result of reading too many PhD comics, but a tinge of anxiety is setting in. What if my advisor turns out to be a workaholic? What if I can't finish in the required 3 1/2 years before my funding runs out? What if my office mates are insane (they're not, I think) or my experiments all fail?
Then I remember that a million students have survived their PhD just fine before me and a million will again. I may not have the prettiest office (in fact, drab is not an inaccurate description), but at least I'm not sharing with 14 other people like my flatmate. My supervisor has only been nice to me, despite the bollocking that he gave his other PhD student last week. And my project, even though it looks daunting from here, will rest on the foundations of my Masters project, meaning that I have a reasonable idea of where to start.
So, the base camp has been established in the foothills of Mt. PhD. Only the future will show if I scale the summit triumphantly, or freeze to death in a crevice somewhere. Boy, that metaphor took a bleak turn, didn't it?
Saturday, August 30, 2008
Research is Easier if You Make It Up
Yes, I know I haven't written in a good few months. In my defence, I have been kept quite busy by the research for my MSc project. Now that it is done, however, I'd like to share a few thoughts on my first real experience with research.
For this first post, I want to talk about what was perhaps the most humbling experience, and that was how tempting it was to cheat.
Like many research projects, my research was beset with problems. There were contradictory results, vague results, results that were the opposite of what we expected, without any indication why this happened. And often, when I got these results, I would think: "Gee, wouldn't it be nice if I could make up the results I wanted instead."
Now before you cast the first stone, let me be very clear: I did not fake any results, nor will I hopefully ever do so. But it got me thinking. How easy would it really be to fake results? For my MSc project, it would have been really easy. We do not have to hand in the code (although it is possible that the markers may ask for the code if they smell a rat, but let's assume for the sake of the argument that the faked results are completely convincing), so I would not even have to write the programs. I knew how the different experiments were supposed to work, so generating some convincing results would have been easy. The only people to see the results are my supervisor and a second marker. Of those two, only my supervisor could possibly spot fake results, because the second marker is not an expert in the field. If I had gotten any fake results past my supervisor, I would basically be home free.
You might be thinking that that's all very well for a Masters project, but surely in real research faked data would be spotted. But would it really? I agree that you would probably have difficulty faking a whole project: You'd be hard-pressed to answer questions from reviewers of your paper, and anyone trying to repeat the experiments would obviously get very different results. But what about just tweaking that one experiment that's poking a hole in your theory? That would again be very easy and would probably not be spotted unless somebody decides to repeat that exact experiment. If somebody later disproves your theory, well, you got a paper out of it, and nobody can really blame you for not spotting the flaw when all of your experiments were confirming the hypothesis.
Cheating can get even more subtle (choosing your experiments, skimping on controls, omitting results) and harder to spot. So the question is, given how easy it would be to cheat, what, other than personal integrity, is keeping scientists honest?
I believe curiosity and ambition are big factors. If you get results that contradict your hypothesis, you don't just say "Aw, crud", you get excited, because there's another problem to solve. Maybe this new problem will lead to an even bigger discovery than the one you were hoping to make. If you just fake the result, you'll probably never do really ground-breaking science. Worse yet, you might set back other scientists who will not pursue their theories because your "results" seem to have disproved them.
There's also training and your research environment. Never underestimate social conformity, which in this case is a good thing. If everybody around you is excited about research, as most scientists will be, you'll find it very hard to be the cheater, even if you're the only one who knows that your results weren't real. You'll want to be just as good as the rest, and if they can deal with contradictory results, then so can you.
Of course, this only applies if people the people in your research environment let you know about the problems they were having. They may be competitive people who feel that talking about struggling with research is equivalent to showing weakness. If that is the case, I recommend reading some of the many excellent blogs from scientists who are not afraid to talk about their research issues.
One thing that is clear is that you cannot just assume that every result that is published is automatically set in stone. If you think you have a better theory, test it, and if necessary repeat an experiment that has already been done. If enough people do that there might actually be a chance of demasking the cheaters. And that would be another great incentive not to cheat in the first place.
For this first post, I want to talk about what was perhaps the most humbling experience, and that was how tempting it was to cheat.
Like many research projects, my research was beset with problems. There were contradictory results, vague results, results that were the opposite of what we expected, without any indication why this happened. And often, when I got these results, I would think: "Gee, wouldn't it be nice if I could make up the results I wanted instead."
Now before you cast the first stone, let me be very clear: I did not fake any results, nor will I hopefully ever do so. But it got me thinking. How easy would it really be to fake results? For my MSc project, it would have been really easy. We do not have to hand in the code (although it is possible that the markers may ask for the code if they smell a rat, but let's assume for the sake of the argument that the faked results are completely convincing), so I would not even have to write the programs. I knew how the different experiments were supposed to work, so generating some convincing results would have been easy. The only people to see the results are my supervisor and a second marker. Of those two, only my supervisor could possibly spot fake results, because the second marker is not an expert in the field. If I had gotten any fake results past my supervisor, I would basically be home free.
You might be thinking that that's all very well for a Masters project, but surely in real research faked data would be spotted. But would it really? I agree that you would probably have difficulty faking a whole project: You'd be hard-pressed to answer questions from reviewers of your paper, and anyone trying to repeat the experiments would obviously get very different results. But what about just tweaking that one experiment that's poking a hole in your theory? That would again be very easy and would probably not be spotted unless somebody decides to repeat that exact experiment. If somebody later disproves your theory, well, you got a paper out of it, and nobody can really blame you for not spotting the flaw when all of your experiments were confirming the hypothesis.
Cheating can get even more subtle (choosing your experiments, skimping on controls, omitting results) and harder to spot. So the question is, given how easy it would be to cheat, what, other than personal integrity, is keeping scientists honest?
I believe curiosity and ambition are big factors. If you get results that contradict your hypothesis, you don't just say "Aw, crud", you get excited, because there's another problem to solve. Maybe this new problem will lead to an even bigger discovery than the one you were hoping to make. If you just fake the result, you'll probably never do really ground-breaking science. Worse yet, you might set back other scientists who will not pursue their theories because your "results" seem to have disproved them.
There's also training and your research environment. Never underestimate social conformity, which in this case is a good thing. If everybody around you is excited about research, as most scientists will be, you'll find it very hard to be the cheater, even if you're the only one who knows that your results weren't real. You'll want to be just as good as the rest, and if they can deal with contradictory results, then so can you.
Of course, this only applies if people the people in your research environment let you know about the problems they were having. They may be competitive people who feel that talking about struggling with research is equivalent to showing weakness. If that is the case, I recommend reading some of the many excellent blogs from scientists who are not afraid to talk about their research issues.
One thing that is clear is that you cannot just assume that every result that is published is automatically set in stone. If you think you have a better theory, test it, and if necessary repeat an experiment that has already been done. If enough people do that there might actually be a chance of demasking the cheaters. And that would be another great incentive not to cheat in the first place.
Saturday, July 5, 2008
Conference Noises
Would you say that a scientists first conference is like his first kiss*, a unique experience, never forgotten despite the fumbling and nervousness? Or is it more like the first time you went to a McDonald's: Sure, it's exciting and colourful, but after you've been a few dozen times you notice that they're all the same.
I couldn't say yet which of these is a better description, since I've only just experienced my first conference. Conference might be saying a bit much: It was a one-day symposium, and I didn't even have to leave the city.
Still, there were some memorable experiences to be made. Some were of the mundane variety: It seems that even in Britain, coffee break means coffee break, and not tea break. And don't even dare ask for water. Also, pinning your badge to your shirt is a fashion faux-pas; the correct place is discreetly on your belt.
The poster session was different from what I expected, because there were really only posters. Somehow, I always expected the poster creators to be standing next to them with proud smiles, eager to explain their science to anyone passing by. Not so here: There were posters, there were people reading the posters, and that was it.
The talks ranged from the fascinating to the mystifying. I've always been better at learning things from papers than at picking them up in lectures, so it's no surprise that I couldn't follow some of the more complicated topics. Listing to those lectures was not a waste of time, though, since at least now I know those topics exist and I can find out more about them (by reading papers!) if I want to.
The quality of the speakers varied (doesn't it always?) but some of them were very good, even inspirational. There are so many unsolved problems in bioinformatics, but these speakers were pointing the way to solving many of them.
Now for the more disappointing part of the symposium. No, not the food, that was alright. This is something that I'm willing to be not many attendees even noticed, but it's actually a huge statistical fluke if it was random: Out of 15 speakers, not a single one was female. I'm used to gender bias in my field, especially on the informatics side, but 0 out of 15? Seriously? You're telling me that there's not a single female professor that you could have invited to talk about her research?
At least many of the people in the audience were female, but jeez!
*With the first conference occasionally preceding the first kiss by a while.
I couldn't say yet which of these is a better description, since I've only just experienced my first conference. Conference might be saying a bit much: It was a one-day symposium, and I didn't even have to leave the city.
Still, there were some memorable experiences to be made. Some were of the mundane variety: It seems that even in Britain, coffee break means coffee break, and not tea break. And don't even dare ask for water. Also, pinning your badge to your shirt is a fashion faux-pas; the correct place is discreetly on your belt.
The poster session was different from what I expected, because there were really only posters. Somehow, I always expected the poster creators to be standing next to them with proud smiles, eager to explain their science to anyone passing by. Not so here: There were posters, there were people reading the posters, and that was it.
The talks ranged from the fascinating to the mystifying. I've always been better at learning things from papers than at picking them up in lectures, so it's no surprise that I couldn't follow some of the more complicated topics. Listing to those lectures was not a waste of time, though, since at least now I know those topics exist and I can find out more about them (by reading papers!) if I want to.
The quality of the speakers varied (doesn't it always?) but some of them were very good, even inspirational. There are so many unsolved problems in bioinformatics, but these speakers were pointing the way to solving many of them.
Now for the more disappointing part of the symposium. No, not the food, that was alright. This is something that I'm willing to be not many attendees even noticed, but it's actually a huge statistical fluke if it was random: Out of 15 speakers, not a single one was female. I'm used to gender bias in my field, especially on the informatics side, but 0 out of 15? Seriously? You're telling me that there's not a single female professor that you could have invited to talk about her research?
At least many of the people in the audience were female, but jeez!
*With the first conference occasionally preceding the first kiss by a while.
Sunday, June 1, 2008
Thursday, May 22, 2008
Fun with Proteins
I think almost anyone who has studied protein 3D structures would agree that it is a hard problem.
For the uninitiated, proteins are made out of chains of amino acids. Each amino-acid consists of a backbone and side-chains. The properties of the side-chains determine the structure that the chain will take on in 3D. For example, polar side-chains may repell each other, and hence tend not to be close. Another example is hydrophobic ("water-fearing") side-chains which need to be on the inside of the protein, away from the water molecules that surround a protein inside the cell.
There's more than one way to determine protein 3D structure. You can take the actual protein, crystallise it, shoot x-rays at it and work out the structure from seeing how the x-rays diverge. Or you can take all of the contraints mentioned above, encode them in a computational model and get a computer to crunch the numbers for you until it finds the optimal structure.
Or, well, you could just get people to do it by hand. For free. And have fun while they're doing it.
That's the principle behind FoldIt, a new game based on, yes, you guessed it, protein structures. The idea is that protein folding is much like a puzzle, and people love doing puzzles. So we let them fold virtual proteins, and evaluate the structures based on the constraints that we know about. Add some fun sounds when you're tugging and dragging proteins, a bonus for reducing the number of moves starting from the initial configuration, and an element of competitiveness in the form of an online ladder, and you've got a fun little game that people will actually want to play.
FoldIt is currently in open Beta and completely free to play. I've tried it out, and it really is a lot of fun. The online option allows you to chat with fellow folders while you're playing, and the interface is simple and intuitive. You don't even have to know anything about proteins; there's a very easy tutorial to get you up to speed, and you'll figure out soon enough what works and what doesn't.
It will be interesting to see if FoldIt players can come up with better protein foldings than a computer could. FoldIt is not the first instance of a "useful" game I've come across. I first heard about the concept, called human computation, in the context of a game called ESP that gets its users to label images with text. It makes you wonder what other arduous bioinformatics tasks we could turn into games (gene-finding anyone?).
For the uninitiated, proteins are made out of chains of amino acids. Each amino-acid consists of a backbone and side-chains. The properties of the side-chains determine the structure that the chain will take on in 3D. For example, polar side-chains may repell each other, and hence tend not to be close. Another example is hydrophobic ("water-fearing") side-chains which need to be on the inside of the protein, away from the water molecules that surround a protein inside the cell.
There's more than one way to determine protein 3D structure. You can take the actual protein, crystallise it, shoot x-rays at it and work out the structure from seeing how the x-rays diverge. Or you can take all of the contraints mentioned above, encode them in a computational model and get a computer to crunch the numbers for you until it finds the optimal structure.
Or, well, you could just get people to do it by hand. For free. And have fun while they're doing it.
That's the principle behind FoldIt, a new game based on, yes, you guessed it, protein structures. The idea is that protein folding is much like a puzzle, and people love doing puzzles. So we let them fold virtual proteins, and evaluate the structures based on the constraints that we know about. Add some fun sounds when you're tugging and dragging proteins, a bonus for reducing the number of moves starting from the initial configuration, and an element of competitiveness in the form of an online ladder, and you've got a fun little game that people will actually want to play.
FoldIt is currently in open Beta and completely free to play. I've tried it out, and it really is a lot of fun. The online option allows you to chat with fellow folders while you're playing, and the interface is simple and intuitive. You don't even have to know anything about proteins; there's a very easy tutorial to get you up to speed, and you'll figure out soon enough what works and what doesn't.
It will be interesting to see if FoldIt players can come up with better protein foldings than a computer could. FoldIt is not the first instance of a "useful" game I've come across. I first heard about the concept, called human computation, in the context of a game called ESP that gets its users to label images with text. It makes you wonder what other arduous bioinformatics tasks we could turn into games (gene-finding anyone?).
Friday, April 18, 2008
A Cautionary Tale
Remember the asteroid that's coming close to hitting the earth in 2036? Most of us probably heard about it at some point or other, made a quick calculation to see if we would be alive then, and forgot about it. If you looked into it a bit further, you found out that NASA only gave it a 1 in 45000 chance of hitting the earth; not nearly enough to worry about.
But then, shock! horror! what if NASA were wrong? People make mistakes. It's not as if they double-checked these results.* And it's not as if the smartest minds of the planet were working for NASA.** Who could we rely on to find out if there are any problems with NASAs calculations? Oh, I know. Let's ask a 13-year old schoolboy. It's good enough for The Times. And what do you know, the whiz kid places the risk at 1 in 450. Definitely in the "we should worry about this" category.
Except NASA never confirmed it, as stated in that article, and has in fact since denied that the boy's calculations are correct. Mark over at Good Math, Bad Math has a simple explanation why his assumptions are flawed.
I don't mean to put down Nico Marquardt, the boy who made the calculations. He obviously put a lot of effort into this, and it must have been pretty good work to convince so many people. Who I do want to put down is the many many papers who simply picked up this story, based on its "newsworthiness", seemingly with no fact checking at all. One phone call to NASA would have cleared it all up. Is this all we can expect from Old Media?
* They did.
** They are.
But then, shock! horror! what if NASA were wrong? People make mistakes. It's not as if they double-checked these results.* And it's not as if the smartest minds of the planet were working for NASA.** Who could we rely on to find out if there are any problems with NASAs calculations? Oh, I know. Let's ask a 13-year old schoolboy. It's good enough for The Times. And what do you know, the whiz kid places the risk at 1 in 450. Definitely in the "we should worry about this" category.
Except NASA never confirmed it, as stated in that article, and has in fact since denied that the boy's calculations are correct. Mark over at Good Math, Bad Math has a simple explanation why his assumptions are flawed.
I don't mean to put down Nico Marquardt, the boy who made the calculations. He obviously put a lot of effort into this, and it must have been pretty good work to convince so many people. Who I do want to put down is the many many papers who simply picked up this story, based on its "newsworthiness", seemingly with no fact checking at all. One phone call to NASA would have cleared it all up. Is this all we can expect from Old Media?
* They did.
** They are.
Subscribe to:
Posts (Atom)
