Photograph: Heinonlein, CC BY-SA 4.0, via Wikimedia Commons
Knowledge Base · Method

What multi-omics means for plant research

A plant is a blueprint, a set of switches, a workshop of proteins and a stream of chemistry, all at once. Reading those layers together, from the same specimen, is what turns a name in a list into knowledge that can be used, and traced.

← Back to the Knowledge Base
Key facts
  • Multi-omics means reading an organism at several molecular layers, genome, transcriptome, proteome and metabolome, and integrating them into one analysis (Hasin et al., 2017).
  • The layers follow the flow of information in the cell: DNA → RNA → protein → metabolite. Each layer is one step closer to what the plant is actually doing.
  • Medicinal compounds live in the metabolome; the instructions for making them live in the genome. Only paired layers from the same plant connect the two (Wolters et al., 2024).
  • Herbarium DNA is fragmented but accumulates almost no miscoding errors; chloroplast genomes have been recovered from specimens up to 146 years old (Staats et al., 2011; Bakker, 2019).
  • 1,656 Madagascar endemics have a preserved specimen only outside the country, and 537 are known from a single specimen worldwide (IsoGentiX dataset).

Ask what a plant is and botany's traditional answer is a name, a description and a pressed specimen in a cabinet. That answer built the science, but it describes the outside of the organism. Since the first plant genome was sequenced in 2000, it has become possible to read the inside as well, and not at one level but at several. Each level has its own "-ome", its own instruments and its own kind of answer. Multi-omics is the discipline of reading them together (Hasin et al., 2017).

The idea is easiest to hold onto if you follow the flow of information inside the cell, because the omics layers map onto it directly. DNA is transcribed into RNA, RNA is translated into protein, and proteins, mostly enzymes, make and transform the small molecules of metabolism. Four layers, one causal chain.

The four layers, one at a time

Genomics: the blueprint

The genome is the complete DNA sequence a plant inherits, every gene and everything between the genes. It is the most stable layer: broadly the same in a root tip and a petal, the same in drought and in rain. Sequencing it answers the question what could this plant do? Somewhere in those sequences sit the genes for every enzyme the species can ever deploy, including, in many medicinal plants, groups of biosynthetic genes that encode the machinery for making defensive chemistry. But a genome alone is a parts list, not a working machine. Most plant genomes contain tens of thousands of genes, and the sequence itself does not say which of them are in use, where, or when.

Transcriptomics: the switches

The transcriptome is the set of RNA messages being copied from the genome at the moment of sampling. Unlike the genome it is dynamic: a leaf under insect attack transcribes a different set of genes from a leaf at rest, and a root transcribes differently from a flower. Reading it, usually by RNA sequencing, answers the question which genes are switched on, here and now? For anyone hunting a biosynthetic pathway this is the layer that narrows the search, because the genes of a pathway tend to be switched on together, in the tissue where the compound is made. Co-expression, genes rising and falling in step, is one of the main clues used to assemble candidate pathways (Wolters et al., 2024).

Proteomics: the machinery

Transcripts are instructions in transit; proteins are the machines actually built. The proteome, read by mass spectrometry, catalogues which proteins are present and in what quantity. It matters because the correlation between transcript and protein is imperfect: messages are made and never translated, proteins persist after their messages fade, and many enzymes are switched on or off after they are built, by chemical modifications the transcriptome cannot see. Proteomics answers which machines are running? For pathway discovery, finding the enzyme itself in the tissue that makes the compound is far stronger evidence than finding its gene.

Metabolomics: the chemistry

The metabolome is the plant's actual chemical output, measured by mass spectrometry or NMR: sugars and amino acids common to all life, and then the specialised metabolites, the alkaloids, terpenoids and phenolics that plants make to defend themselves, attract pollinators and cope with their environments. This layer is the end of the causal chain and, for human uses of plants, the payoff. Morphine, artemisinin, vinblastine and caffeine are all specialised metabolites. When people say a plant "contains" a medicine, the metabolome is where it physically lives. Metabolomics answers the bluntest question of all: what has this plant actually made?

LayerMolecule readMain instrumentQuestion it answers
GenomicsDNADNA sequencingWhat could this plant do?
TranscriptomicsRNARNA sequencingWhich genes are switched on, here and now?
ProteomicsProteinsMass spectrometryWhich machines are actually running?
MetabolomicsSmall moleculesMass spectrometry, NMRWhat chemistry has the plant actually made?

Layer definitions follow Hasin et al. (2017); the pairing of layers for plant specialised metabolism follows Wolters et al. (2024).

Why the layers must come from the same specimen

Here is the catch, and it is the whole argument of this article. Each layer on its own is a description. The explanation lives in the joins between them: this gene, transcribed in this tissue, built this enzyme, which made this compound. Those joins only hold if every layer was read from the same individual plant, sampled at a known place and time.

The reason is biological variation. Two individuals of the same species can differ genetically; two populations in different valleys almost certainly do. Gene expression shifts with tissue, season, weather and stress, and the metabolome shifts with it. In Catharanthus roseus, Madagascar's rosy periwinkle, the route to the anti-cancer alkaloids runs through a long chain of enzymatic steps distributed across different tissues and cell types, which is precisely why the compounds resisted synthesis for decades. Read the genome from one plant and the chemistry from another and you are correlating two different experiments. The causal chain is cut, and no amount of statistics restores it.

Specimen-level multi-omics keeps the chain intact by anchoring everything to one physical object: a well-documented specimen, collected with consent, georeferenced, dated, identified, and vouchered so the identification can be checked by anyone, decades later. From that one plant come the DNA, the RNA, the proteins and the metabolites, and all four data layers point back to the same voucher. The specimen becomes a multi-layered digital object, a record you can interrogate at any level and trust at every level, because the provenance is shared.

From one plant to one linked record
01One documented specimenCollected with consent, georeferenced, dated, identified, and vouchered in an herbarium.
→
02Blueprint and switchesDNA is sequenced for the genome; RNA for the transcriptome, tissue by tissue.
→
03Machinery and chemistryMass spectrometry reads the proteins built and the metabolites actually made.
→
04One linked recordAll four layers attach to the voucher: genotype, expression and chemistry, traceable to one place, date and consent.
Field botany in Madagascar: documenting plants where they grow
Fieldwork is where specimen-level data begins. Everything measured later in a laboratory inherits its meaning, and its legitimacy, from how the plant was documented and consented at collection.Photograph: IsoGentiX field archive
One layer is a description. Four layers from the same plant are an explanation, and an explanation with provenance is one that can be trusted, reused and paid back to its source.

What linked layers make possible

Consider what the joined-up record does that the separate layers cannot. Suppose a metabolomic screen finds an unusual alkaloid in a leaf extract. On its own, that is a spot on a chromatogram. But if the transcriptome of the same leaf is in hand, you can ask which biosynthetic genes were highly expressed in that tissue; if the genome is in hand, you can find those genes' full sequences and their neighbours; if the proteome is in hand, you can confirm which candidate enzymes were actually present. What began as a spot becomes a testable hypothesis about a complete pathway, and pathways are what allow a compound to be studied, produced in cell culture or yeast, and understood, without stripping wild populations of the plant that first made it. This pairing of omics layers is now the standard strategy for decoding plant specialised metabolism (Wolters et al., 2024).

The same logic runs in reverse for conservation. A genome links a species to its relatives and reveals its distinctiveness; expression and chemistry reveal what would be functionally lost if it went extinct. For a flora like Madagascar's, where around 82 percent of some 11,500 native vascular plant species exist nowhere else (Antonelli et al., 2022), almost none of this has been read at any layer. The plants with the most to teach are the least documented, which is the gap IsoGentiX exists to close.

Catharanthus roseus, the rosy periwinkle, in flower
Catharanthus roseus, the rosy periwinkle. The alkaloids that made it one of medicine's most important plants live in its metabolome, invisible to any survey that stops at the species name.Photograph: Biswarup Ganguly, CC BY-SA 3.0, via Wikimedia Commons

The herbarium as a time capsule

Specimen-level thinking also runs backwards in time, because the world already holds hundreds of millions of documented plant specimens: the herbaria. A herbarium sheet is a plant pressed, dried and mounted with its collection data, sometimes centuries ago, and it turns out to be a far better molecular archive than anyone had a right to expect.

The DNA in dried herbarium tissue is heavily fragmented, broken into short pieces by time and preparation. But when researchers sequenced herbarium DNA to characterise the damage, they found that miscoding lesions, the chemical changes that would corrupt the sequence itself, were rare to negligible (Staats et al., 2011). Fragmented but faithful is exactly the kind of damage modern short-read sequencing was built to handle. The field this opened is now called herbarium genomics, or herbariomics: in the work that opened it, chloroplast genome sequences were assembled from dozens of herbarium specimens across twelve flowering-plant families, the oldest of them 146 years old (Bakker, 2019). Later work has even mapped how the fragmentation proceeds over decades, so that sequencing strategies can be tuned to a specimen's age (Weiß et al., 2016).

This matters enormously for a country like Madagascar, because so much of its botanical record exists only as herbarium sheets, and so many of those sheets are overseas. In the IsoGentiX dataset, 1,656 Malagasy endemic species have a preserved specimen somewhere in the world but none in any Madagascar institution, and 537 endemics are known from a single specimen anywhere on Earth. For a species collected once, in a forest that may no longer stand, that one sheet is not just the record of the plant. Increasingly, it is a readable molecular sample of it, the anchor to which any future data layer can be attached.

4molecular layers read from a single specimen: DNA, RNA, protein, metabolite
146years old, the oldest herbarium specimen to yield a chloroplast genome assembly
1,656Madagascar endemics whose only preserved specimen sits outside the country
537endemics known from a single specimen anywhere on Earth

Herbarium sequencing figure: Bakker (2019). Custody figures: computed from the IsoGentiX Flora dataset, reflecting what has so far been digitised and localised.

Ravenala, the traveller's tree, a Madagascar endemic
Ravenala, the traveller's tree, one of Madagascar's most recognisable endemics. Like almost all of the island's flora, its genome, expression and chemistry remain essentially unread.Photograph: phiro, CC BY 4.0, via Wikimedia Commons

Provenance is the ethics

There is a second reason the specimen anchor matters, and it is legal and ethical rather than technical. Under the Convention on Biological Diversity and its Nagoya Protocol, genetic resources belong to the sovereign states they come from. Using them requires the prior informed consent of the country of origin and mutually agreed terms, so that when research on a country's plants produces benefits, a share of those benefits returns to that country (CBD Secretariat, Nagoya Protocol).

That system stands or falls on traceability. Benefit-sharing attaches to the source, and you can only share benefits with a source you can identify. Data that has been cut loose from its origin, a sequence in a database with no collection record behind it, is data whose obligations are unenforceable in practice. Specimen-level multi-omics is, among other things, an answer to that problem: because every layer of data stays attached to one documented, consented collection event, the trail from any future discovery back to the source country, and to the community whose land and knowledge were involved, remains intact. The CARE principles for Indigenous data governance make the same point from the community side: data carries obligations to the people it came from, not just opportunities for those who hold it (Carroll et al., 2020).

Consent comes first

IsoGentiX works by free, prior and informed consent: communities agree to any collection before it happens, understanding what it involves, in their own language and on their own terms. It comes first, always. Knowledge generated from the flora is returned digitally, under national authority, so that the record of a country's plants is held by the country it belongs to.

A note on the legal position. The Nagoya Protocol sets the international framework, but the concrete obligations for any given collection depend on the national access and benefit-sharing law of the country of origin, and how digital sequence information is to be handled is still being negotiated by the parties to the CBD. Nothing on this page states or implies contractual terms; access terms for any IsoGentiX collection are agreed case by case with the national authorities and communities concerned.

Reading as a form of protection

Multi-omics can sound like an argument for extraction: find the compound, take the value. Read at specimen level, under consent, it is closer to the opposite. A flora that has been documented at every layer is a flora whose value is visible, attributable and defensible. The species that makes an unusual molecule becomes harder to dismiss when its forest is weighed against a mine or a plantation. The community that stewarded a medicinal plant has a documented, dated record connecting them to it. And the science can proceed from data rather than from repeated collection of wild plants. That is what the IsoGentiX mission means in practice: we gather the plant world's data, decode it into knowledge, and enable action that protects our environment and serves its people. One specimen, four layers, one traceable record. Decode:Protect.

Common questions

What is the difference between genomics and metabolomics?

Genomics reads the plant's DNA, the full instruction set it inherited, which stays broadly constant across its tissues and its life. Metabolomics measures the small molecules the plant is actually making at the moment of sampling. The genome says what the plant could do; the metabolome says what it is doing, and it is in the metabolome that medicinal compounds physically exist.

Why must all the omics layers come from the same specimen?

Because expression and chemistry vary between individuals, populations, tissues and seasons. A genome from one plant and a metabolome from another cannot be safely joined; the causal chain from gene to transcript to enzyme to compound is broken. Layers read from one documented specimen stay mechanistically linked and traceable to one place, date and consent.

Can old herbarium specimens still be used for genomics?

Yes. Herbarium DNA is fragmented into short pieces, but it accumulates very few miscoding errors, so the sequence that survives is faithful. Chloroplast genome sequences have been assembled from specimens up to 146 years old (Staats et al., 2011; Bakker, 2019).

How does multi-omics relate to benefit-sharing under the Nagoya Protocol?

The Nagoya Protocol ties the use of genetic resources to the consent of the country of origin and to agreed terms for sharing benefits, and that only works if data can be traced to its source. Specimen-level multi-omics keeps every data layer attached to one documented, consented collection event, so any future discovery can be traced back to the country and community it came from.

Sources and further reading

  1. Hasin, Y., Seldin, M. & Lusis, A. (2017). Multi-omics approaches to disease. Genome Biology, 18, 83. genomebiology.biomedcentral.com, the standard introduction to omics layers and their integration.
  2. Wolters, F.C., Del Pup, E., Singh, K.S. et al. (2024). Pairing omics to decode the diversity of plant specialized metabolism. Current Opinion in Plant Biology, 82, 102657. sciencedirect.com, how paired omics layers are used to discover biosynthetic pathways.
  3. Bakker, F.T. (2019). Herbarium genomics: plant archival DNA explored. In Lindqvist, C. & Rajora, O.P. (eds), Paleogenomics: Genome-Scale Analysis of Ancient DNA, 205–224. Springer. link.springer.com, reviews plastome assembly from herbarium specimens up to 146 years old.
  4. Staats, M. et al. (2011). DNA damage in plant herbarium tissue. PLOS ONE, 6(12), e28448. journals.plos.org, herbarium DNA is fragmented but miscoding damage is limited.
  5. Weiß, C.L. et al. (2016). Temporal patterns of damage and decay kinetics of DNA retrieved from plant herbarium specimens. Royal Society Open Science, 3, 160239. royalsocietypublishing.org
  6. Duffin, J. (2000). Poisoning the spindle: serendipity and discovery of the anti-tumor properties of the Vinca alkaloids. Canadian Bulletin of Medical History, 17(1–2), 155–192. pubmed.ncbi.nlm.nih.gov, the history of vinblastine and vincristine.
  7. Antonelli, A. et al. (2022). Madagascar's extraordinary biodiversity: evolution, distribution, and use. Science, 378(6623), eabf0869. science.org
  8. Carroll, S.R. et al. (2020). The CARE Principles for Indigenous Data Governance. Data Science Journal, 19(1), 43. datascience.codata.org
  9. Secretariat of the Convention on Biological Diversity. The Nagoya Protocol on Access and Benefit-sharing. cbd.int/abs, the treaty text and official guidance.