Friday, November 7, 2008
On Learning When To Shut Up
Example:
Me: This is not giving the result I expected. I observed effect Y, so maybe it's due to Z.
Person X: Oh, so you're saying that [slight rephrasing of Z] is causing the problem.
No, I'm not stating an absolute, as indicated by my use of the word "maybe". I'm putting forward a hypothesis, which is what you do in science. But I don't appreciate you nailing me down on that hypothesis before I have even investigated it.
Maybe (just a hypothesis!) the problem is with me. These are (for the most part) busy people I'm talking to, and maybe they can't afford to spend that much time speculating anymore. So when I'm coming up with a hypothesis, they just assume I wouldn't be telling them if I hadn't already thought about it for some time.
In other words, do I need to learn when to shut up?
Monday, October 13, 2008
A Plea
It's not fun to get a matrix without any column or row labels. Like my old physics teacher used to say when somebody proudly proclaimed that the answer to an exercise was 5.029, "5.029 what? Elephants?". In physics, a quantity is not meaningful without a unit, and in data management, a matrix is not useful without labels.
It's not fun to know nothing about any preprocessing of the data. Has it been normalised? I guess I can check the mean and standard deviation, but what if it's only close to 0 and 1, but not exactly so? Was that some special normalisation method? Hell if I know.
It's not fun to know nothing about the experimental setup. Maybe you told me every second data point is a wildtype. Does that mean that these were results from two-colour arrays? Or two singe-colour arrays? Are the wildtypes from the same time point as the mutants? Come to think of it, are these even time-course data?
It's not fun to know nothing about what the biologists* want you to find. Are they looking for similarly expressed genes or for regulators? Would they rather have a network of the knocked-out genes or of all genes? Is it worse if I give them false positives or false negatives?
So please, please, please, document your data. Maybe some time in the future I will tell you about fun things called wikis and databases, but for now even a text file would do.
*Or insert other applied science here.
Saturday, August 30, 2008
Research is Easier if You Make It Up
For this first post, I want to talk about what was perhaps the most humbling experience, and that was how tempting it was to cheat.
Like many research projects, my research was beset with problems. There were contradictory results, vague results, results that were the opposite of what we expected, without any indication why this happened. And often, when I got these results, I would think: "Gee, wouldn't it be nice if I could make up the results I wanted instead."
Now before you cast the first stone, let me be very clear: I did not fake any results, nor will I hopefully ever do so. But it got me thinking. How easy would it really be to fake results? For my MSc project, it would have been really easy. We do not have to hand in the code (although it is possible that the markers may ask for the code if they smell a rat, but let's assume for the sake of the argument that the faked results are completely convincing), so I would not even have to write the programs. I knew how the different experiments were supposed to work, so generating some convincing results would have been easy. The only people to see the results are my supervisor and a second marker. Of those two, only my supervisor could possibly spot fake results, because the second marker is not an expert in the field. If I had gotten any fake results past my supervisor, I would basically be home free.
You might be thinking that that's all very well for a Masters project, but surely in real research faked data would be spotted. But would it really? I agree that you would probably have difficulty faking a whole project: You'd be hard-pressed to answer questions from reviewers of your paper, and anyone trying to repeat the experiments would obviously get very different results. But what about just tweaking that one experiment that's poking a hole in your theory? That would again be very easy and would probably not be spotted unless somebody decides to repeat that exact experiment. If somebody later disproves your theory, well, you got a paper out of it, and nobody can really blame you for not spotting the flaw when all of your experiments were confirming the hypothesis.
Cheating can get even more subtle (choosing your experiments, skimping on controls, omitting results) and harder to spot. So the question is, given how easy it would be to cheat, what, other than personal integrity, is keeping scientists honest?
I believe curiosity and ambition are big factors. If you get results that contradict your hypothesis, you don't just say "Aw, crud", you get excited, because there's another problem to solve. Maybe this new problem will lead to an even bigger discovery than the one you were hoping to make. If you just fake the result, you'll probably never do really ground-breaking science. Worse yet, you might set back other scientists who will not pursue their theories because your "results" seem to have disproved them.
There's also training and your research environment. Never underestimate social conformity, which in this case is a good thing. If everybody around you is excited about research, as most scientists will be, you'll find it very hard to be the cheater, even if you're the only one who knows that your results weren't real. You'll want to be just as good as the rest, and if they can deal with contradictory results, then so can you.
Of course, this only applies if people the people in your research environment let you know about the problems they were having. They may be competitive people who feel that talking about struggling with research is equivalent to showing weakness. If that is the case, I recommend reading some of the many excellent blogs from scientists who are not afraid to talk about their research issues.
One thing that is clear is that you cannot just assume that every result that is published is automatically set in stone. If you think you have a better theory, test it, and if necessary repeat an experiment that has already been done. If enough people do that there might actually be a chance of demasking the cheaters. And that would be another great incentive not to cheat in the first place.
Saturday, July 5, 2008
Conference Noises
I couldn't say yet which of these is a better description, since I've only just experienced my first conference. Conference might be saying a bit much: It was a one-day symposium, and I didn't even have to leave the city.
Still, there were some memorable experiences to be made. Some were of the mundane variety: It seems that even in Britain, coffee break means coffee break, and not tea break. And don't even dare ask for water. Also, pinning your badge to your shirt is a fashion faux-pas; the correct place is discreetly on your belt.
The poster session was different from what I expected, because there were really only posters. Somehow, I always expected the poster creators to be standing next to them with proud smiles, eager to explain their science to anyone passing by. Not so here: There were posters, there were people reading the posters, and that was it.
The talks ranged from the fascinating to the mystifying. I've always been better at learning things from papers than at picking them up in lectures, so it's no surprise that I couldn't follow some of the more complicated topics. Listing to those lectures was not a waste of time, though, since at least now I know those topics exist and I can find out more about them (by reading papers!) if I want to.
The quality of the speakers varied (doesn't it always?) but some of them were very good, even inspirational. There are so many unsolved problems in bioinformatics, but these speakers were pointing the way to solving many of them.
Now for the more disappointing part of the symposium. No, not the food, that was alright. This is something that I'm willing to be not many attendees even noticed, but it's actually a huge statistical fluke if it was random: Out of 15 speakers, not a single one was female. I'm used to gender bias in my field, especially on the informatics side, but 0 out of 15? Seriously? You're telling me that there's not a single female professor that you could have invited to talk about her research?
At least many of the people in the audience were female, but jeez!
*With the first conference occasionally preceding the first kiss by a while.
Sunday, June 1, 2008
Friday, April 18, 2008
A Cautionary Tale
But then, shock! horror! what if NASA were wrong? People make mistakes. It's not as if they double-checked these results.* And it's not as if the smartest minds of the planet were working for NASA.** Who could we rely on to find out if there are any problems with NASAs calculations? Oh, I know. Let's ask a 13-year old schoolboy. It's good enough for The Times. And what do you know, the whiz kid places the risk at 1 in 450. Definitely in the "we should worry about this" category.
Except NASA never confirmed it, as stated in that article, and has in fact since denied that the boy's calculations are correct. Mark over at Good Math, Bad Math has a simple explanation why his assumptions are flawed.
I don't mean to put down Nico Marquardt, the boy who made the calculations. He obviously put a lot of effort into this, and it must have been pretty good work to convince so many people. Who I do want to put down is the many many papers who simply picked up this story, based on its "newsworthiness", seemingly with no fact checking at all. One phone call to NASA would have cleared it all up. Is this all we can expect from Old Media?
* They did.
** They are.
Saturday, March 29, 2008
Open Reviews?
Not that I disagree, necessarily, but I've heard it all so often that it's not really registering anymore. Maybe the killer apps will come along, and maybe they won't. Web 2.0 thinking has potential for science publications, but not everything that has potential gets realised.
But towards the end of the article, there was something that caught my attention: Open review.
Certainly rating a paper would seem reasonable when done by the Faculty of 1000 (http://www.f1000biology.com), but it is not a generally accepted practice. We challenge you to rate this Editorial too. In some ways the reluctance to rate a scientific paper is strange since we suspect the same person may well rate a book on amazon.com. Another option would be to add a Digg or del.icio.us button (http://digg.com or http://del.icio.us) to incorporate conventional media ranking tools into an academic journal Web site. If one finds an interesting article, one could immediately flag it with these tools.Now this is interesting, because peer-review is a nearly sacred notion in science. Your paper has not proven merit until it has passed the peer-reviewing process and been published. It basically says, "other scientists thought this was worth reading".
So what happens if the reviewing suddenly become open to all? Well, it essentially become a popularity contest. Digg is a perfect example: The stories that end up on the front page are the ones that a lot of people liked. Very democratic, isn't it? Only it means that today the stories included "Pranks to pull on your Co-workers" and "The 10 Most Mismatched Movie Couples".
Oh, there were plenty of interesting stories as well, but my point is that what's popular is not necessarily what's best. By all means, let people comment and review papers, but make sure we know which reviews come from scientists, and which come from your average schmoe. I know it sounds elitists, but it's not much more elitist than demanding that the person who takes out your appendix have a medical degree.
Friday, March 21, 2008
Some Reading for your Easter Weekend
FemaleScienceProfessorReading FSP makes me wish that I'd had somebody like her as a lecturer in my undergraduate years, rather than a series of boring white guys. (Not you Dr. S! Nor you Professor W!) She clearly cares about her students and loves her job, always a winning combination. Plus, I find her stories about careless misogyny in academia endlessly fascinating (in a horrifying, hope-I-never-act-like-that way).
Bioinformatics Zen
Not updated very often (kind of like this blog, huh?), but when it is, the articles are always worth reading if you're interested in the nitty-gritty of bioinformatics. Anything related to the field can come up here, whether it's the intricacies of programming, tips on how to get a PhD or humorous characterisations of stereotypical bioinformatics people.
Minor Revisions
This blog is a more recent addition to my reading schedule, but a charming one. Katie gives us a glimpse into the life of a biomedical engineering postgrad, and a very personal glimpse at that. I'm always impressed with people who are willing to share their ups and downs on their blog; it's a (skill? trait? strength?) I don't seem to possess. Katie likes having lots of subscribers to her feed, so go subscribe!
Saturday, March 15, 2008
Talking the Talk
Of course, the advice they give applies to presentations given to fellow scientists, with the objective of introducing your work to them. And in that particular scenario, I probably agree with everything they say.
However, what if the aim of the presentation is not to inform, but to educate? In other words, what if you're giving a lecture? This is very topical for me, as I've just finished a course where students were giving presentations on papers, and I've had to do one of the presentations myself. We disregarded most of the rule they came up with. Were we right to do so? Well, let's look at the rules:
- Be able to give the presentation without support of the slides.
That one's a tricky one, because we were explaining a technique. In my part, I was heavily relying on examples to explain what was happening, and those examples were all on the slides. Could I have done it on the blackboard? Probably, but not without taking considerably more time. Still, we did rehearse a few times, so I think we could have brought the point across even without the slides. Overall, this rule holds.
- No outlines on the slides
Now this I can't completely agree with. Sure, giving an outline is slightly superfluous when you're repeating what it says on the slides. But if you're trying to get an unfamiliar topic across to an audience, reinforcement helps. During the presentations by other groups, I often found myself referring back to the slides when I hadn't caught what they were saying. I think outlines have their place in lecture slides.
- The less text the better
Two problems with this one: The first is the point that I just raised that it helps to refer back to the slide if you missed or were confused by what the speaker was saying. The second is that sometimes, the slides are made available to the audience as a study help before or after the talk. They effectively double as lecture notes, and so it is helpful if they contain enough detail so that you can understand them without the help of the speaker.
On the other hand, too much text can indeed be distracting during the presentation. So I'd advocate a compromise solution here: Keep the slides sparse, but provide detailed lecture notes at the end. Unless you're confident that your speaking ability is good enough to allow your audience to follow along easily and take notes while they do.
- Let us see the data
No argument there: Figures should be clear and big enough so that the audience can get a sense of what it is you're trying to show.
Monday, March 3, 2008
And Now For Something Completely Silly
Steve wrote his article tongue-in-cheek, of course, but he bases it on a very real article that appeared in 1985 (!) in the journal Perception: "On the plausibility of superman's x-ray vision" by J.B. Pittenger.There are three basic conditions that a superman x-ray system must meet to be plausible.
1. Transparency:
The rays must be such that all objects but lead are entirely or almost entirely transparent to them. Lead is always entirely opaque to the rays.2. Color:
The rays and processor must result in Superman perceiving the same colors as would an Earthling viewing the scene in ordinary sunlight.3. Exclusivity:
The rays must permit Superman, but not Earthling standing in line with the reflected rays, to see through normally opaque surfaces.
And they wonder why everybody thinks scientists are a bunch of nerds...
Saturday, February 23, 2008
History in our Genes
"The novel finding is the depth of the resolution we've gone to," said National Institutes of Health neurogeneticist Andrew Singleton, co-author of one of the papers in Nature. "This really lets you start moving towards locating individuals geographically. Previously, we've been able to look at the genome and say, 'This part is from Africa, this is from Asia. Now we can look past that and say, 'It's from this part of Africa or Eurasia.'"
Continued Singleton, "We can use these data to look at other areas of the genome that might have been under particular pressure for survival, and go from there to figuring out what the pressure is. One area that was highlighted was the genes responsible for digesting lactose. In countries where there's milk consumption, that one particular haplotype that allows more efficient lactose digestion has arisen."
They've not only been able to identify populations based on the genome alone, but they've also managed to model how humanity spread around the globe. Our history is encoded in our genes. How cool is that?
And all this was done using only the genomes of about a thousand people. Imagine what will be possible once we have even more data. And on the biology side, we will be able to repeat the same analysis for animals or plants. These are truly exciting times.The full article that appeared in Science can be read here.
Thursday, December 6, 2007
Open Source Genetics?
Modularity in computer science has helped unleash crazy amounts of creativity, and new business models derived from user-generated content. Take Google Maps open-API. Or even HTML itself, which allowed users to create graphically sophisticated pages with no real programming knowledge. By putting the hard stuff into a black box and just letting you access what you need to know, user/producers have been able to focus on creating interesting content quickly and easily. What if, in the next decade, the same group of elite users/coders could do the same thing with corn?They might be a little too optimistic, in my opinion. The allure of Open Source (or indeed any kind of hacking) is that anyone can do it. You don't need much initial investment, beyond the computer which you likely already have. Install Linux, get a GNU compiler of your choice, fire up the text editor and you're in business.
Genetic engineering is not like that. Or maybe it's exactly like that, but at a far grander scale. Instead of a computer*, you need a lab: You need pipettes, petri dishes, microscopes, solutions, PCR machines, microarrays, maybe even a gene sequencer. These things don't come cheap.
But let's assume for a moment that you have all of that already. Then you'll still need the things every programmer takes for granted, the libraries or APIs containing shortcuts to all the common tasks that you don't want to design from the ground up. As a genetic engineer, you'll need promoters, restriction enzymes and specialised vectors, each different depending on what you started with.
It is always possible that in the future, genetics labs and components will become as ubiquitous as computers and code libraries. I'm sure that when we were putting punch-cards into basement-sized supercomputers, open source software development seemed as far away as open source genetic engineering seems today. But the transition still took thirty years. I don't think we have to worry about it just yet.
*Or rather, in addition to.
Thursday, November 15, 2007
Reviewing Fun
First of all, most papers can be summarised pretty easily. However, the summary I would come up with almost never matches with the abstract that the authors wrote. I realise that this is a function of their desire to show every aspect of the paper in their abstract, while I would summarise the most important ones (which might be subjective), but I'm still left with the feeling that most abstracts are not reflective of the gyst of the paper.
Secondly, too many papers overuse references. I've read papers where there's two pages of text and three pages of references. What especially ticks me off is when the mentions a topic and then gives five references for it. We don't need five references, we need one good reference. Maybe two if there are two particularly good papers and you can't decide. Five is just overkill.
Thirdly, and finally, I've noticed a distinct lack of detail in some explanations. Now this is something I can understand if you're trying to boil down a paper to two or three pages for publication. But if you're going to gloss over something, at least say that you're doing so. Also, since this is the 21st century, how about providing a link to your webpage where more detailed information can be found?
Monday, October 1, 2007
Attention conservation notice: It's long, and it's about something which makes eyes glaze over even as tempers flare up, and it's not funny at all. Worse yet, more is on the way. You could always read it later, but time spent now is gone forever.The conclusion? IQ may depend much less on your genetic heritage than previously assumed. I find some of his arguments quite convincing.