Thursday, August 27, 2015

Helix: Vaporware or game-changer for cloud-based genomics

Helix

A while back, I wrote a 3-part due diligence [1, 2, 3] on the cloud-based genomics space, focusing on the competitive landscape around Seven Bridges Genomics.

Lets discuss the potential impact of Helix…

It’s not supposed to make money from consumers

The most important thing about Helix is that it is playing the long-game. It does not intended to make money from consumers. My best guess at the Helix strategy is about loss-leading bottom-up disruption with lock-in.

Concept

Basically, they will sequence your exome (or genome) at no cost to the consumer and take a cut of all APP revenue (30% Apple standard?). While the vision is for community developed content, they will need to seed the ecosystem with a few tools to both validate the developer API and generate interest: queue the first batch of high-profile ‘applications collaborations.’

I’m not too concerned about the infrastructure side, as Illumina should have already solved the hard cloud-based problems with BaseSpace and the hard laboratory integration/data management problems via HiSeq X deployments. I suspect this know-how will transfer.

Bottom-up disruption

At $500 in sequencing cost and a 30% app cut, consumers would need to average ~$1500 per person on “What colour will our kid’s eyes be?” and “When will I go bald?”…lulz…I think not.

The healthcare market is notoriously difficult to change; but, get consumers asking about Viagra, and the doctors may follow. That’s the idea behind Helix. Acclimatize the consumer to ‘sharing’ their genome for an insight, and hope it trickles into the clinic one “why can’t you GATTACA” question at a time.

Lock-in

Once consumers are locked-in to the platform, the real money begins: approved genetic tests that will be sold directly to institutions. Unfortunately, without the consumer adoption, Helix just wouldn’t have the clout and social license to operate to show hospitals out of the 1950s.

What other platform would be able to offer ‘free’ sequencing and have a database of (hopefully) millions of users to back it up? Not to mention, a large chunk of that ‘free’ sequencing cost will be ploughed right back into Illumina’s core business as instrument and consumable sales. None of the stand-alone cloud-based genomics shops can deliver that sort of synergy to the bottom-line.

The platform will then bifurcate between the non-approved ‘consumer’ apps (virtually free: check the price of any flashlight app) and the approved ‘institutional’ apps (gravy train).

Genius, if it works.

Competing vision to Oxford Nanopore’s Metrichor

Maybe I over estimate Illumina’s baby killing desire, but I also view Helix as a direct ‘vision’ challenge to Oxford Nanopore’s Metrichor.

Will everyone have a sequencer, like every lab has a PCR machine (Metrichor)?

Will you send-off sample for sequencing as a service, like every lab orders oligos (Helix)?

Will sequencing be ‘things’ focused (Metrichor) or human focused, moving up the value chain from novelty –> clinic (Helix)?

Only time will tell, but I like that both visions are beginning to be articulated.

Conclusion

The Helix announcement could completely upend the ecosystem. But, as always beware the hype and remember that execution reigns.

Thursday, August 13, 2015

Even hackers have epics: Why we need Mel


Mel Kaye. To the uninitiated, a blank stare; within the sept, a folk hero. Mel’s [Free verse, Prose, Gist; Explained] is an epic tale of ‘real’ programming, about a level of heavy wizardry that only the very elite may ever approach.

As with all folk heroes, his tale has two sides.
On the one hand, there is the explicit demonstration of:
  • Supreme technical mastery,
  • Personal integrity, and
  • Code as self-expression.
On the other hand, there are elements which some may find subversive:
  • Intrinsic value of the hack,
  • Subversion of authority, and
  • Apathy to the ‘commercial’ value proposition.
Yet, only through this dualism does the story succeed in addressing the ethical questions developers face…

Is it just for programmers to subvert management?
Yes. Out of respect, our protagonist refused to report the cause of the bug.
Respect is the currency of the realm.
- j.ello

Is it just to rail against proprietary (or obfuscated) source?
Yes. Information wants to be free. Code needs to be free.
If programmers deserve to be rewarded for creating innovative programs, by the same token they deserve to be punished if they restrict the use of these programs.
- RMS, see also: GNU Manifesto

Mel encompasses the joys and sorrows of an entire discipline.

Mel does in one short story what technical guidelines and seminars can never achieve.

I say embrace the ethos. I say: What would Mel do?

Grok that.

---
R.I.P. Ed
23 September 1926 – 13 August 2014

Friday, July 3, 2015

Commentary: 50 Smartest Companies

I have to admit, I was a little surprised when the MIT Tech Review had 15 / 50 companies in the Biotech sector, specifically the genomics ecosystem, from the big bad of NGS instrumentation Illumina to relatively small analysis shops like DNAnexus

Is 2015 really the year of genomics biotech?

The gains (infographic) in raw throughput are indeed very impressive, but I’m skeptical:

Analysis/interpretation >> Point-of-care/Portable >> “High-throughput”

2636 genomes! 100k genomes! 1M genomes!

Maybe China will throw its hat in the ring and spring for 1B genomes. Maybe folks will get serious about somatic variation and spring for 1k genomes from an individual.

IMHO, it seems like an over-hyped pissing contest of who can pay Illumina (sequencing), AWS (compute) and Oracle/IBM (data-centers) the most.

Hopefully, I’ll be proven wrong, tax payers may rejoice in money not flushed and the Tech Review can be vindicated.

Thursday, June 25, 2015

Recap: PyDataUK 2015

This weekend, ~200 delegates trudged through typical London weather (rain) to the Bloomberg offices in London to attend PyDataUK 2015.
While it’s not your typical nerds in T-shirts meet-up; if you use Python to hack data, this conference is probably definitely for you.

Attendance

Curiously, for a ‘data science’ conference the attendance list (which I would have crawled LinkedIn with…heh), was not available. Bases on my (biased) observations, the attendance was roughly as follows…
Type Sub-type Percentage (%)
Industry 70
Self-employed 20
SME 40
Sponsors 10
Large <1
Academia 30
Ugrad <1
Masters <1
PhD 15
Postdoc 5
Professor 10
Government <1
A few highlights…
* Self-employed contractors and consultants were very well represented.

Conference feel

A your data conference, not a ‘big data’ conference

Hadoop has delivered value for <10% of the companies that have installed it
- Paraphrase, anon
This conference is data focused, i.e. focused on using the Python ecosystem to solve your data challenges. The focus is on practice, and practical tools, not theory.
Type Approx Size Appropriate tools
Micro-data <1Gb Ipython
Small-data (Memory-limited) ~10Gb Pandas
Medium-data (Disk-limited) <1Tb Ad-hoc databases
Big-data Tb - Pb Consider enterprise solutions, or grep
The fact is, ‘big data tools’ would be wildly inappropriate for the vast majority of attendees. The problem seems particularly acute in the life sciences. In his war story talk, Paul Agapow covered the herculean efforts required to re-purpose an ill-advised ‘big data’ solution to recover data from a an ongoing clinical trial.
His message was very clear. Life sciences tends to have very detailed, very heterogeneous data in hundreds to thousands of rows (small/medium data): let the data guide the solutions: you probably don’t need enterprise software, so just don’t waste your money.

A Python is useful conference, not a “Python is deity” conference

All tools are shyte, but some tools (Python!) are useful.
- Paraphrase, anon
Speakers like Russel Winder and his talk on the lack of computation efficiency in Python, even using libraries like numpy set a memento mori undertone to some of the more blatant Python triumphalism.

An interpersonal conference, not a Cloister

The very high-level of interpersonal interaction is yet another way in which the conference betrays the nerds in T-shirts. This is very much a conference that one goes to seek guidance and solve problems.
While there are always the stragglers that don’t head down the pub, a good 2/3s of the conference went for fruitful discussion and drink on Saturday. Unsurprisingly, pub attendance was lower on Sunday, but still fruitful.

A place to get hired/take action, not heavy on theory

Folks were hiring like crazy, and it was very much a sellers market.
If you’re a job seeker anywhere on the Python+data spectrum, I’d strongly recommend attending. Companies were recruiting along the entire spectrum, everywhere from AWS-ineering to user-focused commercial data analysis with IPython notebooks (or re-dash, see Arik’s talk for more details on this user-friendly database interaction framework).
In-line with the action oriented nature of the conference, the Pivigo Recruitment founds were there, doing resume/CV screens and offering advice, both to students and established professionals.
If you are a PhD/Postdoc looking to make the transition, I highty recommend taking a look at their Science to Data Science training program.
Continuum may also be prototyping a training programme of their own through its Client Facing Consultant position. Not entirely sure, but 6-months of training via a 3rd-party consultuncy followed by an intentional poach (Continuum –> 3rd party) could be an interesting model.

Talks

I found the spread of talks fantastic. At least amongst the talks I attended…
Type Percentage (%)
Tools 40
War story 30
Skills 20
Under the hood 10

Tools

Tools talks were the most common. They covered ‘non-brand name’ and upcoming tools with emerging communities.
Attend/watch if:
(i) You want to learn about specific tools that may be applicable to your problem.
(ii) You want to collaborate on extending / adopting new tools.

War story

These talks gave the horrifying and nitty-grity details of a specific problem the speaker faced, and how they went about solving it (including gotcha’s and failures). The focus isn’t ‘wow, look at me’; but rather, this was some B.S., and I want no one to go through what I went through ever again.
  • Paul Agapow: Don’t use ‘big data’ tools when simpler solutions will do, particularly in the life sciences.
Attend/watch if:
(i) You want help with the problems you are immediately facing
(ii) You want exposure to problems you’ve never thought-of.

Skills

These were high-level talks that focused more on skills and knowledge than specific tools.
  • Ian Ozdvald: Writing code for you is only the begining, lets see what it takes to push a Bloomberg model to production.
Attend/watch if:
(i) You want to learn what you need to know in a new area.
(ii) You want an overview of a topic you’ve never heard of.
(iii) You want to chat with the speaker about specific War Stories, after the talk.

Under the hood

These talks focused on low-level implementation details of numpy, pandas, Cython, Numba, etc with a particular focus on performance and appropriateness. Personally, I found these talks the most useful. Where else can one gather such concentrated information from the mouth of the open-source contributors.
  • Russel Winder: If you want performance, use Python as a glue-language, and write your computationally intensive functions in a ‘real’ language.
  • Jeff Reback: In pandas, think about idioms and built-in vectorization to get the most out of your code (then write in a ‘real’ language if you still need to go faster).
  • James Powell: Why does writing good numpy feel so different than writing good Python: because the styles have diverged, and will probably continue to do so.
Attend/watch if:
(i) You want a fire-hose of information about low-level topics.
(ii) You want to know how the ‘magic’ happens.

Take-home

This is very much a conference focused on solutions. If you have a problem, don’t be shy!. Ask around, and there will be people there that have faced similar problems, eager to help.
As for me, I look forward to attending next year!

Thursday, March 12, 2015

Book Review: How to lie with statistics


A data analysts bible for communicating stats to non-experts. A recommended re-read as annual absolution for your statistical sins.

There is no surprise it's a classic: the book has aged remarkably well, the (humourus) anticdoes being as pertinent today as 60 years ago.

The premise is quite straight-forward. When presented with stats, keep in mind:
1) Tools of the trade,
2) Lies, and
3) Fallacies.
Then do a "sniff-test."

Tools are the trade include bias, sample size and significance tests.

Lies are (often graphical) ways of misleading the reader (intentionally for the data scientist; plausibly unintentionally for those with less of a background): changing the scale bars, 'cleverly' chosen percentages, dishonest before/after and my personal favorite, semi-attached figures (what the medical profession now calls 'surrogate end points').

Fallacies include the ever present correlation is of course causation, and 'proving' the null hypothesis.

If a breezy 124 pages is too much, cut straight to the end. At a 'lengthy' (by this books standards) 15 pages, the 10th and final chapter enumerates a 5-step 'sniff test' that can stop a good many lie in its tracks:
1) Who says so?
2) How does he know?
3) What's missing?
4) Did somebody change the subject?
5) Does it make sense (particularly for extrapolations)

If pointy haired boss ever read this book, it'd make the data analysts job -- appease power by bending truth -- 456.7% more challenging!

PS. Speaking truth to power will get you fired 654.3% faster than appeasement. Exercise minimally bent truth with caution. You've been warned!

Sunday, January 18, 2015

The cross-functional team: Separation of concerns

Working on a cross-functional team is hard!

As specialists, be that specialization in software development, bioinformatics, or molecular biology, we are domain experts; yet, projects still fail to come to fruition on time, on budget and with the expected impact. This is as frustrating and demotivating to the non-technical manager as it is to the specialist team members.

Lets look at a case study…

Kate (software developer) and Darnell (biologist) are hustled into a meeting room by Xue (non-technical manager). In good faith, Darnell (biologist) lays bare his frustrations with the existing software. Kate (developer) records these as a list of requirements. After rubber-stamp approval by Xue (non-technical manager) and two weeks of furious coding, the revised software is ready. Unfortunately for Darnell (biologist), the software is even worse than before.

There is another meeting, with more senior developers and biologists in attendance Kate steps-through the changes she made, and how the changes address the requirements gathered from Darnell. The biologist and developers don’t understand much of each-others technical jargon, but do their best to provide input into Kate’s new requirements. The developers insist that ABI Instruments are used by the biologists because their output format is standardized, and therefore easier to import. The biologist demand that ‘big data’ capabilities are implemented by the developers. All Xue can think about is justifying the expense and deadline slippage to her annoyed higher-ups, along with the sickly feeling of having her neck being breathed down.

Sound familiar?

The organizational problem is that our specialists, Kate and Darnell, are trained to deliver ‘locally’ optimal solutions within their area of expertise; unfortunately, real-world problems are usually ‘global’. The challenge to the cross-functional team is to approximate a reasonable ‘global’ solution with a set of ‘local’ solutions contributed by each specialist. Put another way…”software problems” are few; “problems benefiting from software” are many.

Separation of concerns to the rescue

Separation of concerns (SoC) is a precept of modern software engineering. Focusing on the ‘what’, and abstracting the ‘how’, enables collaborative development on large code-bases, easing maintenance, extension and debugging.

The concept is simple: disparate modules of code must communicate through a common interface. As long as the interface remains intact, the internal workings of each individual module may be modified independently. Importantly, each developer need only know the details of their own module, and the interfaces of the modules they interact with. The concerns (implementation level details) of each module, are thus separated (self-contained, and preferably free-standing).

At first glance, this might not seem particularly relevant to Kate and Darnell, but managing separation of concerns should be a cross-functional team’s #1 tactical priority, second-only to sharing a common vision (#1 strategic priority). Separation of concerns forces our cross-functional team to focus on ‘what’, instead of ‘how’.

For the software engineer, separation of concerns means crafting sensible code modules, with a thoughtful API. For the cross-functional team member, it means understanding the high-level problem (what), coming to a common understanding (interface) with ones peers, making that understanding explicit (‘human API’), and sharing a common language to discuss solutions.

To successfully implement, each team member requires:1

Mutual ownership of the ‘global’ solution.
Gradient awareness of team member capabilities.
Commitment to communication, which includes trust, honesty and good faith.

Re-examining the case study…

Kate (software developer) and Darnell (biologist) are hustled into a meeting room by Xue (non-technical manager). In good faith, Darnell (biologist) lays bare his frusterations with the existing software. Kate (developer) politely stops Darnell (biologist), and asks Darnell and Xue about the actual ‘problem’ they’re trying to solve (mutual ownership). Putting aside specific frustrations with the software, the three discuss each others overall-process and ‘pain-points’, both biological and software (gradient awareness). The three state and adjust their understanding, and break to assess (communication):

Kate (developer): Problems that can be solved with software, pro/con for various options?
Darnell (biologist): ditto for molecular biology.
Xue (non-technical manager): Context. What are other teams doing?

After a few days of assessment, there is another meeting. Xue outlines the high-level overview discussed previously to make sure everyone is on the same-page about the problem. Kate (developer) presents the pros/cons of a few software options. Darnell (biologist) follows suit for molecular biology (communication). The group discusses the various options (mutual ownership); in consultation, Xue chooses the set of options to be implemented, and leaves the implementation details to Kate and Darnell. As the week progresses, software and biology changes are applied. Kate and Darnell touching base to reassess their understanding if an interface becomes unclear, or additional dependencies arise (gradient awareness). Their ‘local’ solutions each contribute to solving the ‘global’ problem. Concerns have been separated such that developers aren’t telling biologists how to do their jobs, and vice-versa.

Applying separation of concerns is an art

There are no right answers. The central challenge lies with each specialist coming to a common understanding of their interfaces with peers, walking the tight-rope of openness (about what) and abstraction (about how). Thus, applying separation of concerns simultaneously requires teams to understand more of the overall problem (what), so as to define sensible interfaces between team members, yet less of each-others implementation-level details (how).2.

If the proper balance of openness and abstraction is not achieved, then interfaces are either drawn too broadly (you’ll be stepping on each others toes and exposed to unnecessary implementation-level details), or too narrowly (the ‘local’ solutions of each team-member won’t work together to address the ‘global’ problem).

Done well, separation of concerns enables specialists to work together in harmony: delivering a reasonable set of ‘local’ solutions to the ‘global’ problem at hand.
Done poorly, separation of concerns stifles innovation by imposing artificial barriers: ‘local’ solutions which are ‘globally’ ineffective (and sad panda for all parties).

Cross-functional teams of the world, try giving the principle of separation of concerns a try on your next project.


  1. As corollaries, “not my problem”, dismissal of your peers capabilities and inter-specialty rivalry are unacceptable.
  2. An added-benefit is that less field-specific jargon tends to be used because implementation-level details (how) are abstracted into higher-level problems (what). For example, everyone can understand that a software application is slow (what), but the biologist could (usually) care-less that it’s due to excessive network traffic (why), or that the developer resolved the problem by local caching (how).

Sunday, December 14, 2014

Expat Travel Insurance (Part 2)

If you’ve settled upon travel insurance to cover illness/accident on your visit back to your country of citizenship (queue USA chant), it’s decision time on the exact policy.

There is a staggering array of choice, divided into three main product types:

  • Single-trip
  • Multi-trip
  • Backpacker

While there are tools to narrow the search, I am not aware of any that will take the unique concerns of expats (e.g. cover in country of citizenship) into consideration. This means reading terms and conditions: you cannot take for granted that the policy will be right for you, since the screens available are designed for vanilla residence==citizenship types.

For the UK-based among us, this means using a comparison engine like Money Super Market or Confused.

Comparison engine evilness

  • For underwriting purposes, your age is fair game, but not personal identifying information such as Name, email or phone number. I used the traditional ‘abc xyz’ and abc.xyz@yahoo.com as work arounds.
  • Check the insurance agent directly. You will almost certainly find price discrepancies between going direct with the agent and going through the comparison engine.
  • For visits longer than ~7days, check both single-trip and annual multi-trip rates. I found the break-even point is between the two types of products 7-10 days.

Debenhams evilness

After reading through 1/2 dozen terms and conditions, I settled upon an annual multi-trip policy fronted by Debenhams. Cheaper than many single-trip policies for my length-of-stay.

There is much evilness to be had here, but some of their competitors were even worse…
* £30.60 for the Gold policy from Confused versus the (as far as I can tell identical) £61.51 Superior policy ‘direct’ from Debenhams.
* ‘Direct’ is a bit of a misnomer since Debenhams is a front. The policies are actually underwritten by Rock Insurance, which is itself the UK-front for a Swedish firm named SOLID Försäkringar. If it’s any consolation, both firms are registered with the Financial Services Authority [1, 2]. Though, if anything goes amiss, I have more faith in Debenhams attempting to salvage the reputation of its personal finance business than the FSA.
* The 3-star Defaqto rating is sufficient. Self-insure against luggage loss and cancellation, so the extra coverage that the 4/5 star policies cover in these areas is irrelevant. Given that the legal coverage for ‘high street’ policies is a joke (£15k), all they’re worth is the medical (£10M) and personal liability (£2M) cover.
* If you’re an expat considering travel insurance, be aware that you aren’t really protected against a very serious illness/accident, see Part 1.

Debenhams BEYOND evil

Saving the most evil for last…
* Buried in the terms and conditions there is an automatic renewal clause, BEYOND evil IMHO. Opt-out can be performed online: http://ww2.rockinsurance.com/autorenew/
* Making Melkor look like Wayne Brady, some companies have a written-only or phone-only (invariably an 0845 or other premium rate number) opt-out. Unspeakably evil.

As with all things, YMMV.

Written with StackEdit.

Expat Travel Insurance (Part 1)

At times, US policy seems designed to make life difficult for expats. Look no-farther than the catch 22 of insurance for trips back to the good ol’ US of A, where the protections of European-style socialized medicine do not apply.

At first glance, there are three alternatives:

  • US health insurance
  • Expat health insurance
  • Travel insurance

Each with their own unique pros/cons…

Type Eligibility US treatment Price
US Min 6mo US residence + $$
Expat Min 6mo foreign residence + $$$
Travel Min 6mo residence - $

US Health Insurance

Long-time foreign resident are wholly ineligible for US-based health insurance due to residency requirements.

For those returning stateside, good luck finding an affordable short-term PPACA compliant policy to cover the 6 month gap before eligibility in US plans.

On the upside, bless the wonks, there is an Obama-care exemption for those US citizens which pass either the bona fide resident or physical presence test under USC 26 §5000A(f)(4). Not that most non-executives could afford the premiums demanded by reasonable individual plan.

Expat health insurance

These plans generally cover treatment in either your country of citizenship or country of residence. In a perfect world, this would be the plan-type of choice for expats. Country of citizenship treatment for serious accident/illness; country of residence for more minor issues.

Unfortunately, expat health plans are often even more expensive than equivalent US-based health insurance plans. If you’re a non-exec, it’s doubtful that your company will offer this perk, so good luck affording a policy.

Travel insurance

These come in two flavors, US-based plans and foreign-based plans. Generally, there is a 6 month residence requirement, which will determine whether whether you are eligible for the US-based (min 6mo US residence) or foreign-based (min 6mo foreign residence) plans.

The biggest down-side of travel insurance is that they only cover emergency treatment, i.e. to minimize their cost, they will repatriate you ASAP to your country of residence (usually where the plan is acquired).

For relatively minor injuries (e.g. a broken leg) this level of coverage should be sufficient. However, in the event of a serious accident/illness, this money-saving tactic could kill you. For example, say you were involved in a serious car accident while visiting family in your country of citizenship. Travel insurance would cover your stabilization and medical repatriation to your country of residence (expat home). You’d then be left to the state system (e.g. NHS) with limited social/family support during recovery. Moreover, if you’re unable to work, or otherwise violate your visa conditions due to your illness/accident, you very well may face deported back to your country of citizenship, where you’d be uninsured, and more or less left to die without adequate treatment.

Seriously, it’s that dire. In a place like the US, the best an uninsured former expat could hope for is medical bankruptcy (certainty) and surviving without treatment long enough to become eligible for Medicaid (highly uncertain).

Additionally, some plans have additional restrictions on travel to the ‘home country.’ For example, if you’re an Indian national, resident in the US, your US-based plan may not cover you in India. As of yet, I’ve found no way to screen for these exceptions, other than to (attempt) to read through the 20-100 page terms and conditions for each policy.

Conclusion

None of the above is professional advice…far from it. If you’re in the same unfortunate boat, thrust into the complex tax and compliance situation that US policy imposes on its expat community, but can’t afford appropriate advice, good luck!

“There is no hint that help will come from elsewhere to save you.”
- R.I.P. Sagan.

PS. Given my income constraints, I’ve decided to opt for travel insurance, i.e. adequate cover for a minor illness/accident; wholly screwed in the event of a major problem.

Written with StackEdit.

Tuesday, December 2, 2014

Building blast+ databases with taxonomy ID (taxid_map)

Building NCBI BLAST+ databases with linked taxonomy is far more difficult than it should be.
For example, in taxonomy-based tools such as Kraken, mapping
1) taxonomy id to sequence id (gi or accession) and
2) taxonomy id to a human-readable taxonomy tree,
are built-in and transparent to the user.
Unfortunately, with BLAST+ these steps must be completed manually and are included in two separate programs, makeblastdb for (1) and blastn/blastp/blastx for (2).

(1) Taxonomy id <–> sequence id

In BLAST+, a taxid_map file file must be created and passed to makeblastdb
makeblastdb -in <FASTA file> -dbtype nucl -parse_seqids -taxid_map taxid_map.txt 
where taxid_map.txt is a space or tab separated list of sequence ids (either gi or accession) and taxonomy ids.
For example, with gi:
taxid_map.txt
556927176 4570
556926995 4573
501594995 3914
Alternatively with accession:
taxid_map.txt
NC_022714.1 4570
NC_022666.1 4573
NC_021092.1 3914
There is no turn-key way to generate this mapping taxid to sequence_id for a moderately large set of sequencing.
Fortunately, there is always a hack work-around. NCBI allows export of both the FASTA and GenBank files. The former are used as the default input for makeblastdb, and the latter contain both the sequence_id and taxid. They can both be obtained from the NCBI, searching and exporting with Send to:
enter image description here
This simple Python code snippet will do the trick for small and moderately large datasets.
from Bio import SeqIO
genbankfile = "DNA.gb"
f = open('taxid_map.txt','w')
for gb in SeqIO.parse(genbankfile,"gb"):
    try:
        annotations = gb.annotations['gi']
        taxid = gb.features[0].qualifiers['db_xref'][0].split(':')[1]
        f.write("{} {}\n".format(annotations, taxid))
    except:
        pass
f.close()
For large datasets, the bandwidth cost of of downloading the GenBank from NCBI becomes prohibitive, and the dictionary approach would probably be warranted.
Download both the FASTA and GenBank, alternatively extract the FASTA from GenBank, e.g. with BioPython.

(2) Taxonomy id <–> Taxonomy tree

Simply include this NCBI database in the same directory as your database for the look-up to work with blastn/blastp/blastx: ftp://ftp.ncbi.nlm.nih.gov/blast/db/taxdb.tar.gz

Eating your cake

blastn -db <DATABASE> -query <QUERY> -outfmt "10 qseqid sseqid pident staxids sscinames scomnames sblastnames sskingdoms"
1,gi|312233363|ref|NC_014692.1|,86.26,310261,Sus scrofa taiwanensis,Sus scrofa taiwanensis,even-toed ungulates,Eukaryota
1,gi|223976078|ref|NC_012095.1|,86.26,9825,Sus scrofa domesticus,domestic pig,even-toed ungulates,Eukaryota
1,gi|5835862|ref|NC_000845.1|,86.26,9823,Sus scrofa,pig,even-toed ungulates,Eukaryota
Your comma separated file (-outfmt) showing human-readable taxonomy info.

Tuesday, May 20, 2014

Due Diligence: Seven Bridges Genomics (Part 3)


Drawing
At last, following a genomics industry overview (Part 1) and cloud-based genomics analysis platform macro-view (Part 2), we arrive at the micro-level market landscape surrounding Seven Bridges Genomics.

In Part 2, I mentioned that the cloud-based genomics analysis space is crowded.

To give you a sense of just how crowded (read: very crowded), I’ve enumerated the companies in Seven Bridges Genomics’ competitive sphere. I used my best judgement with respect to direct competitors.

Suffice to say, even when restricted to direct competitors, it really is crowded. There is no clear market leader as of yet, so the next few years are going to very exciting/scary for many of these folks. So many ways to die.

Speaking of ways to be disintermediated before wide-spread platform acceptance, I want to give a special shout-out the Next-gen sequencing companies directly building cloud application to plug-into their own systems: GenapSys, Ion Torrent Systems, Oxford Nanopore.

Oxford Nanopore, still in quasi-shadow-mode as they are, like to be secretive; however, they mention AWS cloud applications in their Early Access Documentation. That, along with the very end-to-end nature and slow-to-market release, leads me to believe it’s in the pipe. I’ll be exploring these guys a bit more, later.

In Part 4, we’ll take a deep-dive into Seven Bridges Genomics, assessing positioning and key risks. Until then, enjoy (aside: sorry about the side-scroll, until I find a better solution for narrow blogger page widths)…
Company 1,2 Location Founding Tags 3,4 Products Cloud Direct Competitor
Agile Genomics* Mt Pleasant, SC 2007? Consulting AlignShop, MiST Database X
Aridhia Informatics* Edinburgh, UK 2008 Healthcare/Clinical AnalytiXagility X
Appistry* St. Louis, MO 2001 Consulting Ayrris X
Ayasdi* San Francisco Bay Area, CA 2008 General Machine Learning Ayasdi Cure/Topological Data Analysis (TDA) X
BGI EasyGenomics* Greater Boston, MA 1999/2010 Nonprofit, Core Facility, Open Source Various X
Bina Technologies* San Francisco Bay Area, CA 2011 Hardware/IT Bina Applications X X
BioDatomics* Greater Washington, DC 2012 Open Source SaaS, Pro, Community5 X X
Congenica* Cambridge, UK 2013 Healthcare/Clinical, Healthcare/Diagnostic Sapienta ? ? (Not Released)
Cypher Genomics Greater San Diego Area, CA 2011 Mantis X X (Early Access)
DNAnexus* San Francisco Bay Area, CA 2009 DNAnexus Platform X X
Eagle Genomics* Cambridge, UK 2008 Consulting ElasticAP X
Era7 Bioinformatics* Granada, Spain; Greater Boston, MA 2004 Consulting, Open Source, Bacterial N/A
Fios Genomics* Edinburgh, UK 2008 Consulting N/A
GenapSys6 San Francisco Bay Area, CA 2010 Hardware/Sequencing Genius X ?
Genestack* Cambridge, UK; St. Petersburg, Russia 2012 Genestack Platform X X (Beta)
Genome Cloud Seoul, Korea ? g-Insight X X
Genomics Limited Oxford, UK 2014 Shadow-mode N/A
GenoSpace Greater Boston, MA 2011 Shadow-mode ? X
Geospiza PerkinElmer* Seattle, WA 1997 Desktop GeneSifter
Globus Genomics Chicago, IL ? Globus Platform X ?
Ion Torrent Systems by Life Technolgies 7* San Francisco Bay Area, CA 2007 Hardware/Sequencing Ion Reporter X ?
Maverixbio* San Francisco Bay Area, CA 2012 Desktop Maverix Analytic Platform
NextBio by Illumina* San Francisco Bay Area, CA 2004 Desktop? NextBio Platform
NZGL Dunedin, New Zealand ? Consulting
Omicia Biocomputing* San Francisco Bay Area, CA 2009 Healthcare/Clinical Opal X
Oxford Gene Technology* Oxford, UK 1995 Desktop, Sequencing Service CytoSure Interpret
Oxford Nanopore 8* Oxford, UK 2005 Hardware/Sequencing X ?
Personalis* San Francisco Bay Area, CA 2011 Consulting, CRO, Sequencing Service N/A
Seven Bridges Genomics* Greater Boston, MA; Belgrade, Serbia (IT) 2009 Igor X X
Spiral Genetics* Seattle, WA 2012? Consulting, Desktop N/A / Anchored Assembly Method
Station X* San Francisco Bay Area, CA 2010 Desktop Gene Pool
Syapse* San Francisco Bay Area, CA 2009 Healthcare/Clinical Synapse Platform X
The Genome Analysis Centre (TGAC)* Norwich, UK 2009 Nonprofit, Core Facility Various X
Tute Genomics* Salt Lake City, UT 2012 Tute Platform/ANNOVAR X ?
Technical Notes:
After a misguided regression/sojourn into typing Part 2 in Google docs then copying to Blogger, which resulted in ultra-crap formating, I’m back to using stacked.io, which I explored here. If there is demand/interest, I’m willing to update/convert this listing to a more dynamic format. Just give a shout in the comments.

  1. To the best of my knowledge, these companies form a closed set under the LinkedIn feature ‘People Also Viewed’, omitting spurious hits.
  2. * Direct link to company LinkedIn Page
  3. Companies are for-profit unless otherwise stated, e.g. Nonprofit
  4. Core facility implies Sequencing Service.
  5. BioDT Community is free to use
  6. Special shout-out for Hardware/Sequencing companies with cloud applications.
  7. Special shout-out for Hardware/Sequencing companies with cloud applications.
  8. Special shout-out for Hardware/Sequencing companies with cloud applications.

Monday, May 19, 2014

Due Diligence: Seven Bridges Genomics (Part 2)


Continuing with the top-down analysis from Part 1, lets look at the cloud genomics analysis industry with a focus on macro-scale phenomena. Seven Bridges Genomics, or any other individual firm for that matter, won't be able to do much about these, other than role with the punches.

In a future post, I'll go micro and drill-down to the unique selling points, enduring competitive advantages and economic moat that make Seven Bridges Genomics value proposition durable and secure (aside: hopefully).

Value Proposition

Provide the tools for scientist to do analysis without having to worry about the details of (1) compute/IT and (2) standardized work-streams.



Those are some hefty assumptions...
(1) Assumes scientists are compute limited.
(2a) Assumes there is a value-add in standardized work-streams
(2b) Which then, in-turn assumes, that there exists standard work-streams.


Making Money

As a private company, I can't do a deep-dive into their financials (aside: woe!), so I have to make some assumptions. From the marketing there seem to be two potential revenue streams...
(A) 'Compute Spread,' basically an interest rate spread but for AWS CPUs. The justify their mark-up over AWS compute pricing based on the perception of value-added. Note that this is a subclass of software as a service.
(B) Consulting

(A) must necessarily dwarf (B). Traditional consulting doesn't scale, which dooms a tech company before it can gain its sea legs / line-cross / other nautical right of passage, i.e. shark VC money. Consulting firms can bootstrap, but that doesn't seem like the growth trajectory they're going for.

 So for simplicity, lets reduce to (A). Taking the spread comes with both top-line and bottom-line risk. 

 Paramount amount them, the bottom-line risk of becoming an AWS whipping boy. You can scream for mercy, not that it helps. Honestly, other than try and take the compute in-house or trade masters.

In-house: Manage to do it even comparable to AWS...ha! 
Trade Masters: High switching cost...if it comes to this, were doomed a long time ago.

On the top-line, they need to either work in an highly inefficient market (alas, big banks) or continuously justify the spread they take through value-add. As I mentioned in the previous post, there is loads of competition with no clear market leader. Market structure will not save them, so value-add they must maintain, less open source eats their lunch.



Macro Swallow


The internet meme of near-misses between whales and humans, including such precious lines as “You’re gunna have to do more than clean that wet suit bro” [Youtube] are the impetus behind this section.

It IS a big ocean; however, there are lots of fishies $£€ to be had in a quite restricted space, the wind-up to a feeding frenzy. There are many ways to die.

Last post I based my mental model on drivers and constraints, but this time around a framework based on relative growth rates seems more suitable. A swallow, in this context, means the facet of growth that trumps the others.

Data swallow

Fail: Value Proposition 1


I/O swallow:
Problem: Impractical to upload data to cloud.
Solution: Co-locate with sequencing centers; however, this requires a) consolidation in sequencing industry (mass-market) or b) working with and servicing big co's exclusively.
Prognosis: Not great. a) Is survivable, but may kill the growth curve. b) Basically become just another IT integrator / service provider. Not scalable. Both mean having an additional whipping masters (AWS + core/big co). 


Storage swallow:
Problem: Impracticable to store data.
Solution: Stream data to be processed in real-time.
Prognosis: Would actually be a boon for Seven Bridges if they could solve the streaming and real-time analysis, as it enhances the value proposition.

Compute swallow

Fail: Tech swings against you

But personal processing power grows even faster:
Problem: New algos or technology lower the compute burden, making the cloud unnecessary. Can go back to on-laptop analysis, where other established firms, e.g. Acelrys, may well eat your lunch.
Solution: Go toe-to-toe away from the cloud. Convince that cloud is worthwhile for other reasons (hassle free, a la Google Docs).
Prognosis: If desktop, grim (infrastructure re-boot). If cloud, fine.

But processing doesn't grow fast enough:
Problem: Can't make money off the AWS spread because tasks are sucking too much compute
Solution: Hope parallelism and clever algo saves you, otherwise...
Prognosis: If AWS can’t do it, neither can you most likely. Hosed.

People swallow

Fail: Value proposition 2

Problem (2a): Scientist don't value your workstreams.
Solution: Hope your compute value proposition holds.
Prognosis: If your API doesn't suck, they build there own in your sandbox IF the compute justification is strong enough. Will become niche for low-end / small-time users, as more sophisticated users disintermediate you and take their algo straight to compute.

Problem (2b): Model fails since there are no standardized workflows. Everything must be custom/application specific.
Solution: Turn into a consulting company.
Prognosis: No scale. Either turn niche, or eaten by a bigger consulting fish with scale in consulting.


Takehome

If you're placing a positive bet on the cloud genome analysis industry, not just Seven Bridges Genomics in particular, you're taking a few implicit assumptions...
  1. I/O swallow will not kill the industry in the cradle.
  2. Compute challenges are Goldilocks.
  3. Bioinformatics is amenable to automation and cross-application standardization.
I'm fairly confident of an all-clear on (2) and (3), but (1) worries me. There are solutions here if the company can pivot fast enough, but I'm not convinced that a start-up, as opposed to a core/big co, has the leverage to pull it off. 

There is also a get out of jail free card...alternate value propositions. 

One that sits quite well for the cloud is integration between different datasets, a task made much easier once all this disperate data is sitting on servers you control. One can imagine mining other peoples data and selling insights. This is NOT consulting in the traditional sense, but scalable returns from data integration and automated analysis.

Only time will tell...

Due Diligence: Seven Bridges Genomics (Part 1)

https://www.sbgenomics.com/

"Demonstrate your learning capabilities," how exactly to do that, I wondered. Develop a mental model! 

I've spent the past few hours reading about the field of genomics & next-gen sequencing, with respect to one firm: Seven Bridges Genomics

First I developed a sense of...
  • Promise -- Hard: Routine genomic diagnosis; Harder: Personalized Medicine
  • Problem -- Next-gen sequencing Data  Actionable Results
  • Solution -- ???

After that, I was bit stuck. How can one summarize an entire field with one mental model, one graphic. 

I thought about...
  • Competitive landscape (SWOT)
  • BCG Matrix (Definite with ? for most firms)
  • Key players (Companies, People, Locations)

But none of those are quite it. What is the root cause of the Problem. Here is a perfectly, imperfect mental model (all mental models are wrong, but some are useful) that seems to be working for me...

Mental Model for Genomics & Next-gen sequencing landscape

I think it comes down to drivers and constraints. Drivers being those things that push a technology forward, which of course require some metric to track changes in status (italics). Constraints being the rate-limiting resource which most hamper development, also complete with metrics (not-shown, but examples would be number of distinct technologies in the pipeline vs. maturity/ETA, cost per base-pair).

Each aspect of the ecosystem, from sequencing  assembly  analysis, has its own unique set of drivers and constraints. 

Now, is there a rate-limiting step in the ecosystem as a whole?  If so, that's as good a place as any to begin with high-impact solution...
  • Sequencing -- Cost and time is already falling, with a healthy pipeline of new technologies (i.e. not a 'Pfizer').
  • Assembly -- Incremental improvements in Robustness and speed. Throwing more compute (cheap!) at it generally seems to do the trick.
  • Analysis -- More data (sequencing & assembly) don't seem to be resulting in more actionable insight. Ding ding. I think we have a winner.

One thing that may temper going for the rate-limiting step is relative easiness of attacking other problems first. They're all hard problems, so lets stick to our guns and go with rate-limiting.

Which brings us full-circle, back to Seven Bridges Genomics and their solutionIgor, a cloud-based analysis framework.

The software is constraint-oriented, knocking down barriers to compute and people. Let our clever architecture and Amazon Web Services (AWS) take care of the computational scaling. Let our clever bioinformaticians do the heavy-lifting, standardizing workflows for common problems, adapting and scaling existing solutions and maybe even banging out something completely novel.

The result -- time and cost savings due to the experience curve effects, standardization and economies of scale. Awesome, no?

It remains to be seen whether they can compete effectively. It's a crowded space, with no clear market leader; but that's a story for Part 2. Other takes here and here.


PS. I also quite enjoyed the play on the Seven Bridges problem (aside: at least I think it's intentional). Change the graph, e.g. bombing a bridge -- which is more or less what they hope to do with the analytics end of things, -- and you can force a solution.

Friday, May 16, 2014

Blogger posting with Markdown using StackEdit

I recently discovered StackEdit, a tool for writing and previewing Markdown.

Now, I’ve been using Markdown for quite sometime, it’s useful for everything from electronic lab notebooks to taking notes on informational interviews.

Finally, through StackEdit’s built-in post to Blogger feature, I’m able to abandon the clunky default interface, and write posts the way they were meant to be written!

Additional benefits include easy code snippets…

def stackedit():
    print 'I love your product!'

Easy inspirational quotes…

“Perfection is Achieved Not When There Is Nothing More to Add, But When There Is Nothing Left to Take Away”

And easy equations (same syntax as LaTeX, by the way)…

ΔG=12kBTln(Keq)

Needless to say, I’m excited about the switch, and would highly recommend giving StackEdit a try yourself!

Thursday, May 15, 2014

Due Diligence: Tessella


I like to do a bit of public due diligence on companies of interest. Here's a brief example of some highlights from the dossier...

Tessella is a medium-sized technology consulting company that hires a small intake of talented developers each year, many of whom have PhDs. I'm interested in Tessella, and where their past employees have gone. I'd want to use a little LinkedIn based analysis tool for this data-dive [source], but until I get access, I'll have to do it the old fashioned way (read: manually, without a slick API...the humanity!)

Macro trends

At the macro-level, massive turn-over is a negative, but if no one ever leaves, that could be bad too. You'd expect at least a few consultants to fall for a client, and go work for them; it's how ideas spread in a knowledge ecosystem. I'm specifically interested in the Boston-office, so lets compare Boston to total...

Data

Gender assignments are based on name/picture. Some of the gender numbers don't add-up due to the unavailability of this information.

Tessella, All Offices
Past: 276
Current: 193
Current + Past: 36
--> Left company: 240
--> Leaver/Current: 1.2
--> Promotion/Unpromoted: 0.23

Tessella, Boston
Past: 17
Current: 13
   Men: 9
   Women: 3
Current + Past: 2
--> Left company: 15
   Men: 13
   Women: 1
--> Leaver/Current: 1.3
--> Promotion/Unpromoted: 0.18
--> Men/Women: 22/4 = 5.5

No excess turn-over in Boston office

Self-explanatory, but some caveats...
1) Boston office is 10 years old, compared to 30 for the company as a whole; however, LinkedIn has a bias towards more recent events. A priori expect leaver/current to be higher for All offices.
2) The company has grown massively, so most of the people that have ever worked for Tessella have worked there is the past 10 years. Negates (1)
It's a wash; I'd say they're comparable.

Internal promotions are consistent between Boston and the rest of company

0.23 versus 0.18. I'd say these numbers are roughly the same given:
1) Small sample size for Boston
2) LinkedIn quirk. Not everyone listed a promotion as a separate job. I'd only detect a promotion if the person put in a separate entry. I personally know people that don't do this, thus 20 year tenures in their most senior position (for a 40 year old).
3) I wish I had a base-line metric for internal promotion. I'd be interesting to compare Tessella to peers.

White Male dominated

22:4 Men:Women for the Boston office. This is technology consulting after all. No surprises there. Of the male consultants, past and present, all but two are white. Curious how gender and diversity compare to peers. (Vet me LinkedIn!...you don't seem to present gender/racial data to the API, so I'd have to have some fun!)

Micro trends

I'm interested in where people did before Tessella, and where they go afterwards. With access to the LinkedIn API (please give me vetted access!) I could run an analysis on everyone in the company. There are 240 leavers, which is do-able by hand, but 1) I'm lazy (in the good way) and 2) to get a sense, I probably don't need to sample everyone. I'm specifically interested in the Boston office, so that's where I'll start...

Data (with homebrew classification)

Note: I made the bolded function/company classifications to help organize the data. There are of course other-ways to organize. There are only 14 people since I can't access information for the 15th.

Function...
Software Engineering: 5
Informatics/Analysis: 5
   Bioinformatics: 1
   Cheminformatics: 1
   Data Science: 1
   Industry R&D: 1
   Tech Consultant: 1
Project Management: 3
   Project Manager: 2
   Account Executive: 1
Self-employed: 1

Company...
Life Sciences Research: 2
   Broad
   Dana-Farber Cancer Institute
Life Sciences Companies: 4
   Life Technologies
   Novartis (2)
   PerkinElmer
Big Technology: 4
   BBN Technologies (Raytheon R&D)
   IDBS (IT consultancy)
   Microsoft (2)
Small Technology: 3
   Complete (Digital Marketing)
   Extreme Reach (Video ads)
   Tokyo Electron (Semi-conductors)
Finance: 1
   HighVista Strategies (Asset Management)

Some range in functional exits, but they are all technically-aligned.

10/14 exits are day-to-day technical. The 2 project managers work at NIBR (Novartis Institutes for BioMedical Research) and Life Technologies, so I'll assume that they are managing technical projects. The account executive works for IDBS, so I'll assume he's selling and overseeing technical projects. That means substantially ALL of the exits are technical. The biggest jump this group has made was to technical sales and support (account manager). It IS possible to not program day-to-day, but don't expect to stray too far from the technology function.

Some range in industry, but it's really either life sciences or tech.

6/14 in life sciences. 7/14 in technology. It's Boston after-all, I suspect the UK exits would be less life sciences dominated. Within the life sciences and technology silos, there is some diversity between Fortune500, research institutions and SMEs. Some are small, but I wouldn't consider any of them start-ups. I'd be interested in talking to the one outlier in Finance. Don't worry, he's still a techie.


Observation: People Leak on LinkedIn

I wasn't explicitly looking for this, but folks leak loads of information on LinkedIn. If it were possible to scrape and analyse this information, it would be possible to independently audit private company numbers accounts, or infer them if they don't release.



From reading profiles on LinkedIn, I inferred that Tessella will earn £21M for 2013. This is spot on with Tessella's own reporting [source], which (scouts honour) I did not view before the data-dive.


Information I observed from looking at profiles...
  • 2014: 250 people.
  • 2014: 30% PA growth in life sciences. 
  • 2011: His (two) offices (of 8) contribute ¼ of total revenue.
  • 2011, 210 staff. 160 earning, 50 admin.
  • 2011: Consulting is growing 20% PA, ⅕ of revenue. US is 17% of revenue.
  • 2011: Oversaw 37 staff, 4M GBP in revenue. 
  • 2010: Grew office to >20, >2M GBP
  • 2010: 88% of revenues are repeat business

How to estimate  £21M for 2013

£4M from his office, which is ¼ of company = £16M in 2011
- 30% life sciences...emphasizing that this is higher than average. Consulting is 20%, that it's mentioned implies it's higher than average. Let's assume 15% PA growth as a modest guess.
£16M in 2011 * 2 years at 15% = £21M in 2013
Simple example, but with a fire hose of data being automatically parsed, it may be possible to glean much much more.

Other LinkedIn Estimates

3:1 Tooth:Tail
£100k/consultant in revenues


Financials

Overall, the information available in the financials [source], line-up well with LinkedIn. Good to know that folks are bending the truth on their profiles.

There is so much gold in financials. I'm like a kid in a candy shop when I read through them. Not entirely sure it's normal, but I get positively giddy reading annual reports (financial tables first, naturally).

Many SMEs live hand-to-mouth; that's ok for start-ups, where you should be being compensated with a risk premium, but is inexcusable for an established SME. Not the case with Tessella, but it's a good idea to double check.

5:1 Tooth:Tail

(191:40   Billable:Non-billable)
Not sure how this compares to peers. My gut told me the LinkedIn estimate was a bit low. Went and checked. My gut was right. More tooth...woo!

Revenues are £110k/consultant

(21M / 191 consultants from 2013 annual report)
I think this places a natural cap on what a consultant could ever expect to earn. I'm guessing this is ball-park for a tech consultancy shop. Operations houses are pulling in 1.5-2x, while strategy houses are likely billing 3-5x.

Compensation Structure

Ok, you can't get this from the LinkedIn data...

  • £25k     Low
  • £45k     Median (est)
  • £65k     Mean
  • £160k   High
Annual report is £12M in wages (+£600k in pension cost) over 191 consultants. That's £65k/consultant in earnings versus (£21M revenue)  Highest paid director at £160k. Entry-level consultant at £25k. A 6x high-low multiple is actually fairly reasonable in this day and age. Though, the directors did gorge themselves a bit in 2012; suspect this was related to the management buy-out though. If you take the distribution of US household income as a guide, the median is about 2/3 of the mean, giving around £45k median wage.


Conclusion

If you're a tech-head, who wants to become/remain a male (kidding) tech-related-head with a technology or life sciences company, Tessella offers good exits. If you wanted to break into other areas of consultancy, say operations or strategy, or into another non-technical functions, one should probably look elsewhere. Loads of money, a la finance, also look elsewhere. 


To their credit, what I've found is exactly what's written on the tin. Tessella doesn't sell anything more, or anything less, than a solid technical training ground for new hires. Solid pass on the basic DD.