Showing posts with label Science. Show all posts
Showing posts with label Science. Show all posts

Tuesday, May 20, 2014

Due Diligence: Seven Bridges Genomics (Part 3)


Drawing
At last, following a genomics industry overview (Part 1) and cloud-based genomics analysis platform macro-view (Part 2), we arrive at the micro-level market landscape surrounding Seven Bridges Genomics.

In Part 2, I mentioned that the cloud-based genomics analysis space is crowded.

To give you a sense of just how crowded (read: very crowded), I’ve enumerated the companies in Seven Bridges Genomics’ competitive sphere. I used my best judgement with respect to direct competitors.

Suffice to say, even when restricted to direct competitors, it really is crowded. There is no clear market leader as of yet, so the next few years are going to very exciting/scary for many of these folks. So many ways to die.

Speaking of ways to be disintermediated before wide-spread platform acceptance, I want to give a special shout-out the Next-gen sequencing companies directly building cloud application to plug-into their own systems: GenapSys, Ion Torrent Systems, Oxford Nanopore.

Oxford Nanopore, still in quasi-shadow-mode as they are, like to be secretive; however, they mention AWS cloud applications in their Early Access Documentation. That, along with the very end-to-end nature and slow-to-market release, leads me to believe it’s in the pipe. I’ll be exploring these guys a bit more, later.

In Part 4, we’ll take a deep-dive into Seven Bridges Genomics, assessing positioning and key risks. Until then, enjoy (aside: sorry about the side-scroll, until I find a better solution for narrow blogger page widths)…
Company 1,2 Location Founding Tags 3,4 Products Cloud Direct Competitor
Agile Genomics* Mt Pleasant, SC 2007? Consulting AlignShop, MiST Database X
Aridhia Informatics* Edinburgh, UK 2008 Healthcare/Clinical AnalytiXagility X
Appistry* St. Louis, MO 2001 Consulting Ayrris X
Ayasdi* San Francisco Bay Area, CA 2008 General Machine Learning Ayasdi Cure/Topological Data Analysis (TDA) X
BGI EasyGenomics* Greater Boston, MA 1999/2010 Nonprofit, Core Facility, Open Source Various X
Bina Technologies* San Francisco Bay Area, CA 2011 Hardware/IT Bina Applications X X
BioDatomics* Greater Washington, DC 2012 Open Source SaaS, Pro, Community5 X X
Congenica* Cambridge, UK 2013 Healthcare/Clinical, Healthcare/Diagnostic Sapienta ? ? (Not Released)
Cypher Genomics Greater San Diego Area, CA 2011 Mantis X X (Early Access)
DNAnexus* San Francisco Bay Area, CA 2009 DNAnexus Platform X X
Eagle Genomics* Cambridge, UK 2008 Consulting ElasticAP X
Era7 Bioinformatics* Granada, Spain; Greater Boston, MA 2004 Consulting, Open Source, Bacterial N/A
Fios Genomics* Edinburgh, UK 2008 Consulting N/A
GenapSys6 San Francisco Bay Area, CA 2010 Hardware/Sequencing Genius X ?
Genestack* Cambridge, UK; St. Petersburg, Russia 2012 Genestack Platform X X (Beta)
Genome Cloud Seoul, Korea ? g-Insight X X
Genomics Limited Oxford, UK 2014 Shadow-mode N/A
GenoSpace Greater Boston, MA 2011 Shadow-mode ? X
Geospiza PerkinElmer* Seattle, WA 1997 Desktop GeneSifter
Globus Genomics Chicago, IL ? Globus Platform X ?
Ion Torrent Systems by Life Technolgies 7* San Francisco Bay Area, CA 2007 Hardware/Sequencing Ion Reporter X ?
Maverixbio* San Francisco Bay Area, CA 2012 Desktop Maverix Analytic Platform
NextBio by Illumina* San Francisco Bay Area, CA 2004 Desktop? NextBio Platform
NZGL Dunedin, New Zealand ? Consulting
Omicia Biocomputing* San Francisco Bay Area, CA 2009 Healthcare/Clinical Opal X
Oxford Gene Technology* Oxford, UK 1995 Desktop, Sequencing Service CytoSure Interpret
Oxford Nanopore 8* Oxford, UK 2005 Hardware/Sequencing X ?
Personalis* San Francisco Bay Area, CA 2011 Consulting, CRO, Sequencing Service N/A
Seven Bridges Genomics* Greater Boston, MA; Belgrade, Serbia (IT) 2009 Igor X X
Spiral Genetics* Seattle, WA 2012? Consulting, Desktop N/A / Anchored Assembly Method
Station X* San Francisco Bay Area, CA 2010 Desktop Gene Pool
Syapse* San Francisco Bay Area, CA 2009 Healthcare/Clinical Synapse Platform X
The Genome Analysis Centre (TGAC)* Norwich, UK 2009 Nonprofit, Core Facility Various X
Tute Genomics* Salt Lake City, UT 2012 Tute Platform/ANNOVAR X ?
Technical Notes:
After a misguided regression/sojourn into typing Part 2 in Google docs then copying to Blogger, which resulted in ultra-crap formating, I’m back to using stacked.io, which I explored here. If there is demand/interest, I’m willing to update/convert this listing to a more dynamic format. Just give a shout in the comments.

  1. To the best of my knowledge, these companies form a closed set under the LinkedIn feature ‘People Also Viewed’, omitting spurious hits.
  2. * Direct link to company LinkedIn Page
  3. Companies are for-profit unless otherwise stated, e.g. Nonprofit
  4. Core facility implies Sequencing Service.
  5. BioDT Community is free to use
  6. Special shout-out for Hardware/Sequencing companies with cloud applications.
  7. Special shout-out for Hardware/Sequencing companies with cloud applications.
  8. Special shout-out for Hardware/Sequencing companies with cloud applications.

Monday, May 19, 2014

Due Diligence: Seven Bridges Genomics (Part 2)


Continuing with the top-down analysis from Part 1, lets look at the cloud genomics analysis industry with a focus on macro-scale phenomena. Seven Bridges Genomics, or any other individual firm for that matter, won't be able to do much about these, other than role with the punches.

In a future post, I'll go micro and drill-down to the unique selling points, enduring competitive advantages and economic moat that make Seven Bridges Genomics value proposition durable and secure (aside: hopefully).

Value Proposition

Provide the tools for scientist to do analysis without having to worry about the details of (1) compute/IT and (2) standardized work-streams.



Those are some hefty assumptions...
(1) Assumes scientists are compute limited.
(2a) Assumes there is a value-add in standardized work-streams
(2b) Which then, in-turn assumes, that there exists standard work-streams.


Making Money

As a private company, I can't do a deep-dive into their financials (aside: woe!), so I have to make some assumptions. From the marketing there seem to be two potential revenue streams...
(A) 'Compute Spread,' basically an interest rate spread but for AWS CPUs. The justify their mark-up over AWS compute pricing based on the perception of value-added. Note that this is a subclass of software as a service.
(B) Consulting

(A) must necessarily dwarf (B). Traditional consulting doesn't scale, which dooms a tech company before it can gain its sea legs / line-cross / other nautical right of passage, i.e. shark VC money. Consulting firms can bootstrap, but that doesn't seem like the growth trajectory they're going for.

 So for simplicity, lets reduce to (A). Taking the spread comes with both top-line and bottom-line risk. 

 Paramount amount them, the bottom-line risk of becoming an AWS whipping boy. You can scream for mercy, not that it helps. Honestly, other than try and take the compute in-house or trade masters.

In-house: Manage to do it even comparable to AWS...ha! 
Trade Masters: High switching cost...if it comes to this, were doomed a long time ago.

On the top-line, they need to either work in an highly inefficient market (alas, big banks) or continuously justify the spread they take through value-add. As I mentioned in the previous post, there is loads of competition with no clear market leader. Market structure will not save them, so value-add they must maintain, less open source eats their lunch.



Macro Swallow


The internet meme of near-misses between whales and humans, including such precious lines as “You’re gunna have to do more than clean that wet suit bro” [Youtube] are the impetus behind this section.

It IS a big ocean; however, there are lots of fishies $£€ to be had in a quite restricted space, the wind-up to a feeding frenzy. There are many ways to die.

Last post I based my mental model on drivers and constraints, but this time around a framework based on relative growth rates seems more suitable. A swallow, in this context, means the facet of growth that trumps the others.

Data swallow

Fail: Value Proposition 1


I/O swallow:
Problem: Impractical to upload data to cloud.
Solution: Co-locate with sequencing centers; however, this requires a) consolidation in sequencing industry (mass-market) or b) working with and servicing big co's exclusively.
Prognosis: Not great. a) Is survivable, but may kill the growth curve. b) Basically become just another IT integrator / service provider. Not scalable. Both mean having an additional whipping masters (AWS + core/big co). 


Storage swallow:
Problem: Impracticable to store data.
Solution: Stream data to be processed in real-time.
Prognosis: Would actually be a boon for Seven Bridges if they could solve the streaming and real-time analysis, as it enhances the value proposition.

Compute swallow

Fail: Tech swings against you

But personal processing power grows even faster:
Problem: New algos or technology lower the compute burden, making the cloud unnecessary. Can go back to on-laptop analysis, where other established firms, e.g. Acelrys, may well eat your lunch.
Solution: Go toe-to-toe away from the cloud. Convince that cloud is worthwhile for other reasons (hassle free, a la Google Docs).
Prognosis: If desktop, grim (infrastructure re-boot). If cloud, fine.

But processing doesn't grow fast enough:
Problem: Can't make money off the AWS spread because tasks are sucking too much compute
Solution: Hope parallelism and clever algo saves you, otherwise...
Prognosis: If AWS can’t do it, neither can you most likely. Hosed.

People swallow

Fail: Value proposition 2

Problem (2a): Scientist don't value your workstreams.
Solution: Hope your compute value proposition holds.
Prognosis: If your API doesn't suck, they build there own in your sandbox IF the compute justification is strong enough. Will become niche for low-end / small-time users, as more sophisticated users disintermediate you and take their algo straight to compute.

Problem (2b): Model fails since there are no standardized workflows. Everything must be custom/application specific.
Solution: Turn into a consulting company.
Prognosis: No scale. Either turn niche, or eaten by a bigger consulting fish with scale in consulting.


Takehome

If you're placing a positive bet on the cloud genome analysis industry, not just Seven Bridges Genomics in particular, you're taking a few implicit assumptions...
  1. I/O swallow will not kill the industry in the cradle.
  2. Compute challenges are Goldilocks.
  3. Bioinformatics is amenable to automation and cross-application standardization.
I'm fairly confident of an all-clear on (2) and (3), but (1) worries me. There are solutions here if the company can pivot fast enough, but I'm not convinced that a start-up, as opposed to a core/big co, has the leverage to pull it off. 

There is also a get out of jail free card...alternate value propositions. 

One that sits quite well for the cloud is integration between different datasets, a task made much easier once all this disperate data is sitting on servers you control. One can imagine mining other peoples data and selling insights. This is NOT consulting in the traditional sense, but scalable returns from data integration and automated analysis.

Only time will tell...

Due Diligence: Seven Bridges Genomics (Part 1)

https://www.sbgenomics.com/

"Demonstrate your learning capabilities," how exactly to do that, I wondered. Develop a mental model! 

I've spent the past few hours reading about the field of genomics & next-gen sequencing, with respect to one firm: Seven Bridges Genomics

First I developed a sense of...
  • Promise -- Hard: Routine genomic diagnosis; Harder: Personalized Medicine
  • Problem -- Next-gen sequencing Data  Actionable Results
  • Solution -- ???

After that, I was bit stuck. How can one summarize an entire field with one mental model, one graphic. 

I thought about...
  • Competitive landscape (SWOT)
  • BCG Matrix (Definite with ? for most firms)
  • Key players (Companies, People, Locations)

But none of those are quite it. What is the root cause of the Problem. Here is a perfectly, imperfect mental model (all mental models are wrong, but some are useful) that seems to be working for me...

Mental Model for Genomics & Next-gen sequencing landscape

I think it comes down to drivers and constraints. Drivers being those things that push a technology forward, which of course require some metric to track changes in status (italics). Constraints being the rate-limiting resource which most hamper development, also complete with metrics (not-shown, but examples would be number of distinct technologies in the pipeline vs. maturity/ETA, cost per base-pair).

Each aspect of the ecosystem, from sequencing  assembly  analysis, has its own unique set of drivers and constraints. 

Now, is there a rate-limiting step in the ecosystem as a whole?  If so, that's as good a place as any to begin with high-impact solution...
  • Sequencing -- Cost and time is already falling, with a healthy pipeline of new technologies (i.e. not a 'Pfizer').
  • Assembly -- Incremental improvements in Robustness and speed. Throwing more compute (cheap!) at it generally seems to do the trick.
  • Analysis -- More data (sequencing & assembly) don't seem to be resulting in more actionable insight. Ding ding. I think we have a winner.

One thing that may temper going for the rate-limiting step is relative easiness of attacking other problems first. They're all hard problems, so lets stick to our guns and go with rate-limiting.

Which brings us full-circle, back to Seven Bridges Genomics and their solutionIgor, a cloud-based analysis framework.

The software is constraint-oriented, knocking down barriers to compute and people. Let our clever architecture and Amazon Web Services (AWS) take care of the computational scaling. Let our clever bioinformaticians do the heavy-lifting, standardizing workflows for common problems, adapting and scaling existing solutions and maybe even banging out something completely novel.

The result -- time and cost savings due to the experience curve effects, standardization and economies of scale. Awesome, no?

It remains to be seen whether they can compete effectively. It's a crowded space, with no clear market leader; but that's a story for Part 2. Other takes here and here.


PS. I also quite enjoyed the play on the Seven Bridges problem (aside: at least I think it's intentional). Change the graph, e.g. bombing a bridge -- which is more or less what they hope to do with the analytics end of things, -- and you can force a solution.

Saturday, April 19, 2014

Data Dive: Personality testing

Notice: I've been informed that the data were indeed crawled from a combination Adult & Youth surveys. There is no scientific value here, just a funny plug on the power of numpy, matplotlib and a few lines of bash/awk to take a hack at some data.

The Via Institute on Character offers a fun and informative ~10 min personality test. I was curious about how my results compared to the norm, so I took a little data dive.

Here are some humorous observations from a subset of their rich data-set (N=41513)...
  1. People admit to lacking Humility, Self-Regulation and Spirituality.
  2. Honesty, Love and Kindness are people's top priorities.
  3. Spirituality and forgiveness go (very weakly) hand-in-hand.

Observation #1
People admit to lacking Humility, Self-Regulation and Spirituality.

A priori, I expected the character attributes to have flat distributions (my null hypothesis), a straight line at p=0.04 (1/24 attributes). This couldn't be further from the truth for some attributes. 

In the upper-right hand corner of some images, you'll see a (+) or  (-), this corresponds to a big deviation from the flat null hypothesis.
(+) attributes are ranked higher than expected.
(-) attributes are ranked lower than expected. 
Others, without a marker, are more or less flat, e.g. Curiosity and Humor.


Observation #2
Honesty, Love and Kindness are people's top priorities.
Another slightly more flashy way to view the data is through a rank abundance plot. The  top quartile (ranks 1-6) are dominated by these three attributes: 


Observation #3
Spirituality and forgiveness go (very weakly) hand-in-hand.
Correlation coefficient plots by quartile show some very weak (+/- 0.15-0.2) correlation for top ranked (1st Q) and bottom ranked (4th Q) traits, but virtually none for average ranks (2nd and 3rd Q).

Q1 (Top Ranked attributes):
+) Spirituatlity / Forgiveness
+) Humor / Prudence+Perspective
-) Perseverance / Social Intellegence
-) Creativity / Gratitude
-) Curiosity / Perspective
-) Perspective / Creativity+Humility 

Q4 (Bottom ranked attributes):
+) Prospective / Hope+Humor
-) Appreciation / Forgiveness
-) Creativity / Humility

Note that correlation does not imply causality and the correlations, are very weak. They are, however, above the background average correlation of -0.02. These are humorous insights, and in no way should they be taken seriously.

Your Homework
Take the test for youself at viacharacter.org
If you're keen to explore, redacted datasets (taken down at the request of the VIA Institute on Character) and analysis scripts available at https://github.com/rmharrison/viacharacter-analysis



Friday, May 18, 2012

Even Worse Than it Looks: The African-American Male Life Sciences PhD

Summary: Statistics reporting Blacks can be misleading. Be sure to ask about African-Americans (Black US Citizens), in addition to all Blacks. The data are probably even worse than presented.

Given the already low percentage of Black Life Sciences PhD graduates, it's difficult to imagine that the picture could get any bleaker.

It does...

Be sure to look at African-Americans (Black US Citizens), in addition to all Blacks. You may be surprised by the difference such a subtle distinction can make.
..hate to say I told you so.
The NSF also provides a number of pre-built tables, some of which are available here. For example, Table 24 shows the percentage of all Black PhD recipients per field for 2010. Keep in mind that this table shows the results for ALL Blacks. The picture is even worse for African-American males (Black male US Citizens).

...if you're still whirling in disbelief, enter the following into the NSF SED Tabulation Engine:
Outer: Race & Ethnicity (standardized)
Row: Citizenship (survey-specific)
Column: Gender
Filter: Academic Discipline, Broad (standardized)=LIFE SCIENCES