The original interview can be found here
Which plant variety will deliver the best yield in five years – under conditions that may look very different from today? That’s exactly the question Computomics answers with AI models that combine genetics, environmental data and field measurements to assess breeding candidates before the first field trial even takes place. The Tübingen-based AgriTech company has now raised €6.3 million in a Series B round to further establish its platform in commercial breeding programs. In our conversation with founder and CEO Dr. Sebastian J. Schultheiss, we discuss the idea behind Computomics, the technology in detail, and the next steps following the financing.
Plant breeding often takes many years before a new variety reaches the market. What exactly is the problem, and why is the classic breeding process no longer sufficient today?
A breeding program is bound to nature: one generation per season, and at the end a single evaluation that only reflects the current year and location. Depending on the crop, it takes ten to fifteen years for a variety to reach the market. For a long time, that wasn’t a problem, because the target environment barely changed during that period. Today that’s no longer true: if you cross plants now, you’re breeding for a climate that doesn’t exist yet at the time of market launch. On top of that, there’s a scale problem: a crossing population can generate millions of candidates, but only a few thousand can be tested in the field. This selection determines the success of the program, and it’s made based on very few observations. Classical breeding has given us good yields for 10,000 years, but it’s too slow for rapid climatic change.
How exactly does your technology work? What precisely does your AI model predict, and which data feeds into it – genetics, climate, soil?
Our model predicts the performance of a specific breeding candidate in a specific target environment: yield, quality parameters, or stress tolerance, each with an uncertainty estimate. That’s different from the question of which line is “the best” – the answer depends on location, year, and management. To do this, we combine three data layers: the genetics of the candidates and parent lines, increasingly represented as a pangenome rather than a single reference genome; environmental data such as weather, soil, and management practices; and the customer’s field trial data, which calibrates the model. Instead of linear models, we use machine learning, because the interesting effects aren’t additive. For us, the genotype-by-environment interaction isn’t noise – it’s the signal.
How does your AI approach affect the effectiveness and speed of seed development?
First, selection: if you know before sowing which ten percent of candidates have a real chance, you only test those in the field. That saves cycles, because field capacity is the bottleneck, not genetics. Second, precision: we can select for environments that can’t be tested at your own location, such as a drier climate or a different sales market. With AB InBev, we helped shorten barley breeding from twelve to five years. A cycle still takes one season – you just need fewer of them.
You already work with some of the largest seed producers worldwide. How do you know that your predictions make a difference for your customers’ breeding programs in practice?
We have ourselves measured prospectively. The usual setup: we predict performance for the current season, and our prediction is compared against what was measured in the field. What matters isn’t the correlation on known data, but the hit rate in the top decile, because selection happens at the top. The tougher indicator is commercial: breeding programs are conservative, and anyone who extends a method over several years and expands it to more crops has proven an internal impact. That’s publicly documented at AB InBev, which names us as their only external technology partner in their ESG report.
The large seed companies have their own data science teams. Why do they still buy from you, and what does your approach mean for mid-sized breeders?
Because the question isn’t whether you build a model yourself or buy one, but how many iterations you can afford. An internal team builds a model for its own crop and its own data and learns from that once. We see the same problem across crops, climate zones, and programs, and that knowledge sits in the method, not in a single customer’s data. For mid-sized breeders, the math is simpler: hiring a team that masters both genomics and machine learning is more expensive and slower than accessing a platform that already exists.
Breeding data is a breeder’s capital. What happens to the data a customer entrusts to you?
That’s the first question in almost every initial conversation, and it’s a fair one. Our position is clear: customer data is kept separated per customer and deleted on request. Trained models belong to the customer and are not used for work with others. What we reuse is architecture and pipelines – not data, and not models. The second common concern is data volume. Practically every program has enough data to get started. The real work lies in generating the right data going forward, and that’s where we support our customers.
Europe is debating genetic engineering and new breeding techniques intensively. Where does your approach stand in that debate?
Outside of it, which is an advantage. We don’t change the genome at all – we predict which cross is worth making. What ends up in the field is conventionally bred and requires no additional approval. For customers, that means the time gain is immediately usable and doesn’t depend on how regulation looks in five years. Where genome editing is permitted, both approaches complement each other, because even there, the question remains in which genetic background and elite material the new variety is bred.
Congratulations on successfully closing your €6.3 million financing round! What exactly will you use the fresh capital for in the coming months?
First, product: we’re bringing more of the work from our project engagements into the platform so customers can run calculations themselves. Second, team: we’re bringing on more experts who support breeding programs during onboarding. The most common objection in initial conversations is that customers don’t have enough data – that’s almost never true, and the real task is helping customers generate meaningful data going forward. Third, crops and data foundation: we’re expanding our pangenome base to more species and deepening our models where climate directly affects yield.
HTGF has been at your side since the seed stage. How has this collaboration developed over the years, and what has it brought you?
HTGF was our first institutional financing in 2015, at a time when bioinformatics for plant breeding wasn’t a category investors were looking at. Since then, the fund has participated in every round, including now in the Series B. Two things have practically helped. First, access: doors to investors and industry partners that we wouldn’t have had coming out of Tübingen, and regular pointers to support programs we hadn’t been aware of ourselves. Second, reliability: through the Covid years and through our shift from a services business to a product business, we had a shareholder who didn’t get nervous every quarter.
Where will Computomics stand in 2030 – what direction are you taking in terms of crops and markets?
By 2030, genomic prediction should no longer be a large-corporation-only project, but a standard tool, including for mid-sized breeders without their own data science team. We’re not naming a target number of programs; the ambition is breadth: maize, wheat, soybean, sugar beet, and tomato on one side, berries, additional vegetable varieties, forage and ornamental plants on the other. Our vision is that every new variety is developed with our technology. For specialty crops, there’s still almost no infrastructure for data-driven breeding. Geographically, the greatest potential lies in markets where adaptation doesn’t mean optimization, but food security. Our work with IRRI points in that direction.
You’re based in Tübingen, between the Max Planck Institute and Cyber Valley. What do you look for when hiring?
People who speak two languages. Pure machine learning profiles exist, plant geneticists exist – people who bridge both are rare. A model can simulate crosses perfectly and still suggest something that could never actually happen in a real breeding plan. Whoever can see that coming has understood both sides. We mostly train these people ourselves, which is why we recruit heavily from the local ecosystem. On September 9, we’ll be at Cyber Valley Day with a booth in the Startup Basecamp, talking about exactly. After this financing round, we’re hiring for several new positions.
Thank you very much, Sebastian, for your time and insights!
Share on