Sunday, January 27, 2008

Interesting Article, Bad Title

LiveScience had an article about a recent discovery in genetics made by some researchers in Maryland. They found out that DNA sequences that are the same are more likely to cluster together than those that are different.

Curiously, DNA with identical sequences of bases were roughly twice as likely to gather together as DNA molecules with different sequences.

[...]

The electrically charged chains of sugars and phosphates of double helixes of DNA cause the molecules to repel each other. However, identical DNA double helixes have matching curves, meaning they repel each other the least, Leikin explained.
It's an interesting discovery, although it remains to be seen how useful it's going to be. But that's not the reason I'm mentioning it here. What made me laugh was the title they gave the article:
DNA Molecules Display Telepathy-like Quality.

How long until the first New Age healer picks up on this and claims he can heal your DNA by telepathy?

Wednesday, January 23, 2008

Woe is PhD

So I'm looking for a PhD place. Ideally, I'd like it to be at my current university, as the other universities in the UK that are good in my field are mostly located in London, and I don't like that city much. Nor can I really afford to live there.

I'd like it to involve Bioinformatics, but not require any wetlab work that I would have to do myself. There should also be scope for applying Machine Learning techniques. There's two areas of research that interest me. One is work in genetics, such as gene regulation modelling, or protein structure and function prediction. The other is to model biological systems at the macro level and predict how changes in the environment influence animal population size or plant growth.

I could also see myself doing a straight Machine Learning project without any Bioinformatics involved, but with slightly less enthusiasm.

I'd like to have a supportive supervisor, who I can talk to before I start my PhD and who will advise me on how to write up my project proposal. He or she doesn't need to be an academic superstar, but a fair number of puplications and at least some amount of recognition in either Bioinformatics or Machine Learning would not go amiss.

I'd like to get the chance to teach during my PhD, either tutorials or even lectures.

A scholarship would be helpful. If I don't get one, my parents could help, but I'd like to be able to support myself for once.

And tomorrow, I will meet with a potential supervisor who might be able to offer the place that has most of these characteristics. (I'm not sure about the teaching yet.) Fingers crossed!

Friday, December 21, 2007

Google's Newest Invasion

Having cornered the search market, taken over Youtube and struck fear into the hearts of publishers, Google is licking its lips and looking for new targets. Next stop, Wikipedia.

The potential Wikipedia killer app that Google is developing is called knol, and it has of course sparked much discussion in the blogosphere, with bloggers outbidding each other in who can come up with the wittiest pun. (My favourite: Google Sets its Guns on the Grassy Knol). But puns aside, should Wikipedia be worried? Should we?

Yes and no. You have to remember that Google does not always suceed. Remember Google Video? That didn't take off. Google News? I don't know any News junky who uses that. And I don't see people rushing to pick up Google Talk.

Knol operates on a completely different model from Wikipedia. Instead of having a page for a topic that everybody can edit if they think they know better than the original author, in Knol, the author controls the content, and other users can only suggest changes. Also, you could have more than one "knol" on each topic. Google assures us that more popular (and hence, presumably, more accurate) knols will float to the top, but will they really?

Of course, you have a similar problem in Wikipedia, but it's mitigated because you can correct information: In knols, it seems you can only decide if you like the information or not.

Then there is the problem of orphaned topics: What happens if a "good" knol is abandoned by the author? Can we reuse it to start a new knol? Or is that information frozen in time forever, neither to be reused nor updated?

There are tons of problems that could bring knols to its knees. Of course, there are also reasons why it could succeed: There's less risk of vandalism like Wikipedia has seen, and there are more incentives for people to contribute (Google has agreed to share ad revenue if knol owners let them place ads on their pages).

Whether or not knol succeeds, I believe that the two models are different enough that they could even exist side-by-side. After all Encarta and the Encyclopedia Britannica both sold copies, didn't they?

Thursday, December 13, 2007

Pay Per Use Bioinformatics Software

Equinox, a company based in London and started by the Imperial College, has started offering access to some of its bioinformatics tools on a pay per use basis. Basically, they host the software on their servers, and you pay them each time you want to run a query. Their flagship product has to do with protein structure prediction:
The first product available will be Equinox's leading Phyre(TM) homology modelling and fold-recognition software. User research has shown that proteomics is an ideal target market with positive feedback from research, biotech and pharma audiences.
Yadda, yadda, the press release goes on to state how great this all is. Of course my knee-jerk reaction was: "Pay? Nevah!" But then I realised that you'd have to pay anyway. Previously, this product had only been available via a software licence. Now, that's fine for big companies and universities, who use it frequently. I suspect most of them will stick with the licence as well.

But for smaller companies, or an isolated researcher who may only need to use the software once in his career, the Pay Per Use model may actually have an advantage. Of course, it would be preferable if they'd offer the tool for free, but what kind of business model is that?

Unfortunately, I couldn't find out what the actual price per use was, or how it compared to the cost of a licence. How many uses before a licence would be cheaper? They need to get the balance right, otherwise people will opt for licences every time, and the PPU idea might well die a premature death, as far as bioinformatics is concerned.

Meanwhile, if you want to play around with a bare-bones version of Phyre, you can still go here. Just put on your academic hat first.

Thursday, December 6, 2007

Open Source Genetics?

Here's an interesting premise: Wired Blog has an article on "The Open Organism: Genetic Engineering in the Open Source Era". What would happen if you applied the principles of Open Source Software to genetic engineering?
Modularity in computer science has helped unleash crazy amounts of creativity, and new business models derived from user-generated content. Take Google Maps open-API. Or even HTML itself, which allowed users to create graphically sophisticated pages with no real programming knowledge. By putting the hard stuff into a black box and just letting you access what you need to know, user/producers have been able to focus on creating interesting content quickly and easily. What if, in the next decade, the same group of elite users/coders could do the same thing with corn?
They might be a little too optimistic, in my opinion. The allure of Open Source (or indeed any kind of hacking) is that anyone can do it. You don't need much initial investment, beyond the computer which you likely already have. Install Linux, get a GNU compiler of your choice, fire up the text editor and you're in business.

Genetic engineering is not like that. Or maybe it's exactly like that, but at a far grander scale. Instead of a computer*, you need a lab: You need pipettes, petri dishes, microscopes, solutions, PCR machines, microarrays, maybe even a gene sequencer. These things don't come cheap.

But let's assume for a moment that you have all of that already. Then you'll still need the things every programmer takes for granted, the libraries or APIs containing shortcuts to all the common tasks that you don't want to design from the ground up. As a genetic engineer, you'll need promoters, restriction enzymes and specialised vectors, each different depending on what you started with.

It is always possible that in the future, genetics labs and components will become as ubiquitous as computers and code libraries. I'm sure that when we were putting punch-cards into basement-sized supercomputers, open source software development seemed as far away as open source genetic engineering seems today. But the transition still took thirty years. I don't think we have to worry about it just yet.

*Or rather, in addition to.

Thursday, November 15, 2007

Reviewing Fun

One of the required courses for my Masters is a literature review on a topic which we can choose ourselves. So I've been reading lots and lots of papers (on Bayesian networks for modelling gene regulation, in case you want to know), and the more I read, the more I can see certain common themes emerge. Not common themes about the topic, mind you, but just about papers in general.

First of all, most papers can be summarised pretty easily. However, the summary I would come up with almost never matches with the abstract that the authors wrote. I realise that this is a function of their desire to show every aspect of the paper in their abstract, while I would summarise the most important ones (which might be subjective), but I'm still left with the feeling that most abstracts are not reflective of the gyst of the paper.

Secondly, too many papers overuse references. I've read papers where there's two pages of text and three pages of references. What especially ticks me off is when the mentions a topic and then gives five references for it. We don't need five references, we need one good reference. Maybe two if there are two particularly good papers and you can't decide. Five is just overkill.

Thirdly, and finally, I've noticed a distinct lack of detail in some explanations. Now this is something I can understand if you're trying to boil down a paper to two or three pages for publication. But if you're going to gloss over something, at least say that you're doing so. Also, since this is the 21st century, how about providing a link to your webpage where more detailed information can be found?

Monday, November 5, 2007

Undergrad Proves Important Theorem

I read on the Wired blog that a 20-year old engineering student from the University of Birmingham has formally proved that a certain Turing machine model invented by Stephen Wolfram* is the simplest possible model.

Turing machines, for those not up on their theoretical computer science, are simple computing machines that Alan Turing conjectured were capable of calculating any computable function (he didn't say anything about efficency).

The usual model is that of a machine with a ticker tape with a sequence of symbols and a number of states. The machine looks at the tape one symbol at a time, and, depending on the symbol and its current state, decides what to do: Go forward one character, go back one character, overwrite the current character, or change state.

What makes this minimal Turing machine special is that it only has two states and three different symbols (sometimes called colours). The student proved that if you take away a colour or a state, it wouldn't be a Turing machine anymore.

Meanwhile, what important result did I discover during my time as an undergraduate? Oh, that's right, I discovered that you can live on nothing but pizza for a week...

*Yes, the same one who created Mathematica.