[image: :seedling:] *What if the answer isn’t in a database—but in the
connections between databases?*
Join us on *Wednesday, October 7* for a webinar with Ehsan Estaji of the
Umeå Plant Science Centre, Sweden:
[image: :link:] Trapped knowledge in the AI-agent era: how PlantGraph turns
multi-database questions into one traversal you can check
PlantGraph connects plant data across the transcriptome, epigenome,
regulome, and proteome, helping researchers turn complex, multi-database
questions into one connected, traceable query.
As part of the monthly AgBioData Webinar Series
<https://www.agbiodata.org/webinars>, Ehsan will show how this approach can
help AI agents retrieve evidence more reliably—and how we can check whether
AI-generated answers actually stay grounded in that evidence.
Wednesday, October 7
| 1P ET | 12P CT | 11A MT | 10A PT |
Find your local time here
<https://www.timeanddate.com/worldclock/fixedtime.html?msg=AgBioData+Oct+202…>
.
Join Zoom Meeting
https://us06web.zoom.us/j/82038356125?pwd=YVFMRElMdEpHZmtObXFvZlA4QVFXQT09
Meeting ID: 820 3835 6125
Passcode: 160683
Title: Trapped knowledge in the AI-agent era: how PlantGraph turns
multi-database questions into one traversal you can check
Abstract: A real biological question is rarely a record — it is a path.
“Which salt-responsive genes also carry a promoter methylation change, a
transcription-factor binding site, and a known phosphosite?” has no home
database: the answer exists only as a traversal across four of them. This
is why so much of what we collectively know sits trapped. Not because the
facts are missing — they are curated, somewhere — but because the
connections between them were never stored anywhere. Researchers rebuild
those connections by hand, one question at a time, and lose them the moment
they close the tabs.
Capable AI agents raise the cost of that fragmentation rather than
resolving it. An agent handed a set of endpoints has to plan the join
itself, guessing identifiers and directions at every hop, and every hop is
an opportunity to improvise. In our own benchmarking, giving a strong model
tools and API access over fragmented sources did not reliably fix
retrieval; connecting the data first, and retrieving over it
deterministically, did.
PlantGraph (plantgraph.se) integrates community plant resources into one
connected knowledge graph and separates two jobs that agents usually
conflate: the graph retrieves the evidence, the language model only writes
over it. A multi-database question becomes a single traversal — and because
that traversal actually ran, every claim can be followed back to the query
that produced it and the source record that curated it.
I will walk through questions spanning transcriptome, epigenome, regulome
and proteome resolving in one step, how we measure when the generated
narration drifts from the retrieved evidence, and where the approach still
fails. PlantGraph <https://plantgraph.se/> is not a replacement for the
databases it draws on — it stands on them, and every resource that connects
makes the rest worth more.
See also this news blog <https://www.agbiodata.org/node/652>.
Marcela
--
Marcela Karey Tello-Ruiz, PhD
AgBioData Program Manager
https://www.agbiodata.org/
Phoenix Bioinformatics
[image: :seedling:] *What if the answer isn’t in a database—but in the
connections between databases?*
Join us on *Wednesday, October 7* for a webinar with Ehsan Estaji of the
Umeå Plant Science Centre, Sweden:
[image: :link:] Trapped knowledge in the AI-agent era: how PlantGraph turns
multi-database questions into one traversal you can check
PlantGraph connects plant data across the transcriptome, epigenome,
regulome, and proteome, helping researchers turn complex, multi-database
questions into one connected, traceable query.
As part of the monthly AgBioData Webinar Series
<https://www.agbiodata.org/webinars>, Ehsan will show how this approach can
help AI agents retrieve evidence more reliably—and how we can check whether
AI-generated answers actually stay grounded in that evidence.
Wednesday, October 7
| 1P ET | 12P CT | 11A MT | 10A PT |
Find your local time here
<https://www.timeanddate.com/worldclock/fixedtime.html?msg=AgBioData+Oct+202…>
.
Join Zoom Meeting
https://us06web.zoom.us/j/82038356125?pwd=YVFMRElMdEpHZmtObXFvZlA4QVFXQT09
Meeting ID: 820 3835 6125
Passcode: 160683
Title: Trapped knowledge in the AI-agent era: how PlantGraph turns
multi-database questions into one traversal you can check
Abstract: A real biological question is rarely a record — it is a path.
“Which salt-responsive genes also carry a promoter methylation change, a
transcription-factor binding site, and a known phosphosite?” has no home
database: the answer exists only as a traversal across four of them. This
is why so much of what we collectively know sits trapped. Not because the
facts are missing — they are curated, somewhere — but because the
connections between them were never stored anywhere. Researchers rebuild
those connections by hand, one question at a time, and lose them the moment
they close the tabs.
Capable AI agents raise the cost of that fragmentation rather than
resolving it. An agent handed a set of endpoints has to plan the join
itself, guessing identifiers and directions at every hop, and every hop is
an opportunity to improvise. In our own benchmarking, giving a strong model
tools and API access over fragmented sources did not reliably fix
retrieval; connecting the data first, and retrieving over it
deterministically, did.
PlantGraph (plantgraph.se) integrates community plant resources into one
connected knowledge graph and separates two jobs that agents usually
conflate: the graph retrieves the evidence, the language model only writes
over it. A multi-database question becomes a single traversal — and because
that traversal actually ran, every claim can be followed back to the query
that produced it and the source record that curated it.
I will walk through questions spanning transcriptome, epigenome, regulome
and proteome resolving in one step, how we measure when the generated
narration drifts from the retrieved evidence, and where the approach still
fails. PlantGraph <https://plantgraph.se/> is not a replacement for the
databases it draws on — it stands on them, and every resource that connects
makes the rest worth more.
See also this news blog <https://www.agbiodata.org/node/652>.
Marcela
--
Marcela Karey Tello-Ruiz, PhD
AgBioData Program Manager
https://www.agbiodata.org/
Phoenix Bioinformatics
Dear AgBioData Consortium members,
A quick announcement following our call for nominations for a new member of
the AgBioData Steering Committee: *We are extending the nomination deadline
by one week*.
*Nominations are now due Monday, September 28, 2026.*
If you are interested in serving on the Steering Committee, please complete
the *nomination form* <https://forms.gle/7g6JUFJJghxauzNdA>. You may also
nominate a colleague, provided you have confirmed that they are interested
in serving.
We encourage members to take advantage of the additional week to consider
putting themselves forward or nominating someone who would bring valuable
perspectives and expertise to the Steering Committee.
Thank you for your continued engagement with the AgBioData Consortium!
Best,
Marcela K. Tello-Ruiz
On behalf of the AgBioData Steering Committee
--
Marcela Karey Tello-Ruiz, PhD
AgBioData Program Manager
https://www.agbiodata.org/
Phoenix Bioinformatics
*Call for Nominations to the AgBioData Steering Committee*
The AgBioData Consortium is seeking nominations
<https://forms.gle/7g6JUFJJghxauzNdA> for a new member of the Steering
Committee. The Steering Committee is responsible for the governance of the
Consortium and for managing its activities, including our NSF-funded
Research Coordination Network (RCN) grant, *“Reimagining a Sustainable Data
Network to Accelerate Agricultural Research and Discovery.”*
*To be eligible, nominees must be:*
- A member of the AgBioData Consortium;
- Actively involved with agricultural biological data (e.g., curation,
data management, research, software development, library services, etc.);
and
- An enthusiastic advocate of FAIR data principles.
Each Steering Committee member serves a *three-year term* and is expected
to actively participate in the following activities:
- Participate in regular Steering Committee meetings and follow-up
action items;
- Plan monthly AgBioData webinars and discussions;
- Lead consortium-wide white papers and publications;
- Plan and participate in annual all-hands meetings;
- Contribute to funding proposals; and
- Help shape future objectives and directions for the Consortium, in
collaboration with the membership.
*Nominations close September 21, 2026.*
To express your interest in serving on the Steering Committee, please
complete the nomination form <https://forms.gle/7g6JUFJJghxauzNdA>.
To nominate someone else, please provide their name and email address *after
confirming that they are interested in serving*.
Elections will be held in *September/October 2026*.
Marcela K. Tello-Ruiz
On behalf of the AgBioData Steering Committee
--
Marcela Karey Tello-Ruiz, PhD
AgBioData Program Manager
https://www.agbiodata.org/
Phoenix Bioinformatics
Greetings AgBioData Colleagues!
A friendly reminder of the September 2nd AgBioData webinar, we look forward
to seeing you and hope your schedule allows you to join us tomorrow!
===
*From single cells to FAIR data! *
Working with single-cell data? Make your next GEO submission FAIR from the
start! Join us next *Wednesday, September 2nd*, to learn best practices for
submitting single-cell data to NCBI’s Gene Expression Omnibus (GEO).
As part of the monthly AgBioData Webinar Series
<https://www.agbiodata.org/webinars>, *Emily Clough* of the National Center
for Biotechnology Information (NCBI) will present the talk titled:
“*Single-cell
data submissions to NCBI’s Gene Expression Omnibus (GEO)*.” Full abstract
and Zoom link below.
We hope you can join us!
Best,
Marcela
--
*Wednesday, September 2nd, 1PM ET*
| 1P ET | 12P CT | 11A MT | 10A PT |
Find your local time here
<https://www.timeanddate.com/worldclock/fixedtime.html?msg=AgBioData+Sept+20…>
.
Join Zoom Meeting
https://us06web.zoom.us/j/82038356125?pwd=YVFMRElMdEpHZmtObXFvZlA4QVFXQT09
Meeting ID: 820 3835 6125
Passcode: 160683
--
*Speaker: *Emily Clough (NCBI)
*Title:* Single-cell data submissions to GEO
*Abstract: *The Gene Expression Omnibus (GEO,
http://www.ncbi.nlm.nih.gov/geo/) is an international public repository
that archives gene expression and epigenomics data sets generated by
next-generation sequencing and microarray technologies. The GEO repository
is built and maintained by the National Center for Biotechnology
Information (NCBI), a division of the National Library of Medicine (NLM).
For 25 years GEO has been growing and adapting to emerging new technologies
and now contains over 280,000 studies. The newest revolution in
transcriptomics is single-cell technology which accounts for ~30% of all
public RNA-seq studies in GEO. Single-cell data are large, structurally
complex, and present unique challenges for archiving and data delivery. GEO
has updated submission documentation with a section dedicated to
single-cell data (https://www.ncbi.nlm.nih.gov/geo/info/seq.html#singlecell)
describing required data formats, sample organization, and metadata. GEO’s
metadata template provides examples for single-cell RNA-seq and multi-omics
studies.
This webinar will focus on best practices for single-cell submissions to
GEO to meet FAIR (Findable, Accessible, Interoperable, and Reuseable) data
principles.
--
Marcela Karey Tello-Ruiz, PhD
AgBioData Program Manager
https://www.agbiodata.org/
Phoenix Bioinformatics
*From single cells to FAIR data! *
Working with single-cell data? Make your next GEO submission FAIR from the
start! Join us next *Wednesday, September 2nd*, to learn best practices for
submitting single-cell data to NCBI’s Gene Expression Omnibus (GEO).
As part of the monthly AgBioData Webinar Series
<https://www.agbiodata.org/webinars>, *Emily Clough* of the National Center
for Biotechnology Information (NCBI) will present the talk titled:
“*Single-cell
data submissions to NCBI’s Gene Expression Omnibus (GEO)*.” Full abstract
and Zoom link below.
We hope you can join us!
Best,
Marcela
--
* Wednesday, September 2nd, 1PM ET*
| 1P ET | 12P CT | 11A MT | 10A PT |
Find your local time here
<https://www.timeanddate.com/worldclock/fixedtime.html?msg=AgBioData+Sept+20…>
.
Join Zoom Meeting
https://us06web.zoom.us/j/82038356125?pwd=YVFMRElMdEpHZmtObXFvZlA4QVFXQT09
Meeting ID: 820 3835 6125
Passcode: 160683
--
* Speaker: *Emily Clough (NCBI)
*Title:* Single-cell data submissions to GEO
*Abstract: *The Gene Expression Omnibus (GEO,
http://www.ncbi.nlm.nih.gov/geo/) is an international public repository
that archives gene expression and epigenomics data sets generated by
next-generation sequencing and microarray technologies. The GEO repository
is built and maintained by the National Center for Biotechnology
Information (NCBI), a division of the National Library of Medicine (NLM).
For 25 years GEO has been growing and adapting to emerging new technologies
and now contains over 280,000 studies. The newest revolution in
transcriptomics is single-cell technology which accounts for ~30% of all
public RNA-seq studies in GEO. Single-cell data are large, structurally
complex, and present unique challenges for archiving and data delivery. GEO
has updated submission documentation with a section dedicated to
single-cell data (https://www.ncbi.nlm.nih.gov/geo/info/seq.html#singlecell)
describing required data formats, sample organization, and metadata. GEO’s
metadata template provides examples for single-cell RNA-seq and multi-omics
studies.
This webinar will focus on best practices for single-cell submissions to
GEO to meet FAIR (Findable, Accessible, Interoperable, and Reuseable) data
principles.
--
Marcela Karey Tello-Ruiz, PhD
AgBioData Program Manager
https://www.agbiodata.org/
Phoenix Bioinformatics
Hello Everyone,
Join us this *Wednesday, August 5, at 1:00 PM ET* for the next webinar in
the AgBioData Webinar Series <https://www.agbiodata.org/webinars>.
*McKenzie Mabry* (iDigBio), of the *Florida Museum of Natural History,
University of Florida, and the New York Botanical Garden*, will present:
*"Crop Wild Relatives and the Role of Herbaria in Future Food Crop
Security"*
This webinar will explore the important role that herbarium collections and
crop wild relatives play in supporting future food security.
The Zoom link and webinar abstract are included below.
We hope you'll be able to join us!
Best,
Marcela
*Wednesday, August 5th, 1PM ET*
| 1P ET | 12P CT | 11A MT | 10A PT |
Find your local time here
<https://www.timeanddate.com/worldclock/fixedtime.html?msg=AgBioData+Aug+202…>
.
Join Zoom Meeting
https://us06web.zoom.us/j/82038356125?pwd=YVFMRElMdEpHZmtObXFvZlA4QVFXQT09
Meeting ID: 820 3835 6125
Passcode: 160683
--
Speaker:* McKenzie Mabry* (Florida Museum of Natural History, University of
Florida & New York Botanical Garden; iDigBio)
Title: *Crop Wild Relatives And The Role Of Herbaria In Future Food Crop
Security*
Abstract: Although Nikolai Vavilov recognized the potential of crop wild
relatives (CWR) in the early 1900s, the advent of genome editing
technologies such as CRISPR now enables scientists to more fully leverage
CWRs as a source of genetic diversity for cultivated crops. As global
agriculture faces escalating pressures from climate change, plant
biologists are increasingly focused on sustaining crop productivity under
shifting environmental conditions while meeting the demands of a growing
population. CWRs represent a critical reservoir of genetic variation that
can be harnessed to address these challenges. While efforts to expand
germplasm collections of CWRs are ongoing, herbaria remain an underutilized
resource. At the same time, the growing availability of digitized herbarium
data provides new opportunities to address large-scale research questions.
In this study, occurrence records for CWRs of several important crop
species are obtained from repositories such as iDigBio and Global
Biodiversity Information Facility. Following data cleaning, ecological
niche modeling is used to estimate global habitat suitability under both
current and future climate scenarios. This work highlights the increasingly
important role of herbaria and digitized biodiversity data in supporting
crop improvement efforts, particularly in the context of ongoing climate
change.
--
Marcela Karey Tello-Ruiz, PhD
AgBioData Program Manager
https://www.agbiodata.org/
Phoenix Bioinformatics
*Genome editing has revived interest in crop wild relatives as a source of
genetic diversity — but are we fully tapping herbaria to find them? Join us
next Wednesday 5th of August to learn how digitized biodiversity records
and niche modeling are mapping where climate-resilient crop relatives may
be found, now and in the future.*
Hello Everyone,
As part of the monthly AgBioData Webinar Series
<https://www.agbiodata.org/webinars>, McKenzie Mabry (iDigBio) of the Florida
Museum of Natural History, University of Florida & New York Botanical Garden,
will present the talk titled: “*Crop Wild Relatives And The Role Of
Herbaria In Future Food Crop Security*”. Full abstract and zoom link below.
We hope you can join us!
Best,
Marcela
--
* Wednesday, August 5th, 1PM ET*
| 1P ET | 12P CT | 11A MT | 10A PT |
Find your local time here
<https://www.timeanddate.com/worldclock/fixedtime.html?msg=AgBioData+Aug+202…>
.
Join Zoom Meeting
https://us06web.zoom.us/j/82038356125?pwd=YVFMRElMdEpHZmtObXFvZlA4QVFXQT09
Meeting ID: 820 3835 6125
Passcode: 160683
--
Speaker:* McKenzie Mabry* (Florida Museum of Natural History, University of
Florida & New York Botanical Garden; iDigBio)
Title: *Crop Wild Relatives And The Role Of Herbaria In Future Food Crop
Security*
Abstract: Although Nikolai Vavilov recognized the potential of crop wild
relatives (CWR) in the early 1900s, the advent of genome editing
technologies such as CRISPR now enables scientists to more fully leverage
CWRs as a source of genetic diversity for cultivated crops. As global
agriculture faces escalating pressures from climate change, plant
biologists are increasingly focused on sustaining crop productivity under
shifting environmental conditions while meeting the demands of a growing
population. CWRs represent a critical reservoir of genetic variation that
can be harnessed to address these challenges. While efforts to expand
germplasm collections of CWRs are ongoing, herbaria remain an underutilized
resource. At the same time, the growing availability of digitized herbarium
data provides new opportunities to address large-scale research questions.
In this study, occurrence records for CWRs of several important crop
species are obtained from repositories such as iDigBio and Global
Biodiversity Information Facility. Following data cleaning, ecological
niche modeling is used to estimate global habitat suitability under both
current and future climate scenarios. This work highlights the increasingly
important role of herbaria and digitized biodiversity data in supporting
crop improvement efforts, particularly in the context of ongoing climate
change.
--
Marcela Karey Tello-Ruiz, PhD
AgBioData Program Manager
https://www.agbiodata.org/
Phoenix Bioinformatics
Dear AgBioData Community,
*Passionate about teaching, training, or improving agricultural research
data practices? **This one is for you!*
The AgBioData Consortium is launching a new Education & Training Working
Group <https://www.agbiodata.org/working_groups/edu-training>.
Building on the success of the original Education Working Group—which
developed the AgBioData Curriculum for teaching FAIR data practices
<https://zenodo.org/records/14278084>—we are expanding our efforts to
create more accessible, practical, and community-driven training resources
for researchers, educators, students, and database professionals. This next
phase will focus on:
- Identifying AgBioData community training needs and priorities.
- Updating the AgBioData curriculum for diverse audiences.
- Developing a structured AgBioData certification program.
- Creating online training modules on bioinformatics data management,
reporting, publication, and data submission best practices.
Participation requires approximately 2 hours per month for meetings, plus
flexible offline contributions based on your interests and availability.
By joining this working group, you will:
- Help shape the future of AgBioData education and training.
- Earn authorship on publications, training modules, and other community
outputs.
- Contribute your expertise in ways that fit your interests and schedule.
- Connect with a multidisciplinary network of researchers, educators,
curators, and database professionals.
Interested in joining? Sign up here:
https://forms.gle/CLKYAuyrDxqQtiRG7
If you have any questions, please contact Jodi at jhumann(a)wsu.edu.
We hope you’ll join us in building the next generation of FAIR data
education resources for the agricultural genomics community.
Kind regards,
Marcela
On behalf of the AgBioData Steering Committee
--
Marcela Karey Tello-Ruiz, PhD
AgBioData Program Manager
Phoenix Bioinformatics
Hello Everyone,
Join us tomorrow Wednesday, June 3rd at 1 pm ET for an exciting talk
tackling one of plant genomics' most pressing challenges: predicting enzyme
function at scale. Discover how large language models (LLMs), protein
language models (PLMs), and phylogenetics are being combined into an
end-to-end annotation workflow to make it possible.
As part of the AgBioData webinar series, *Gaurav Moghe*, Associate
Professor in the School of Integrative Plant Science at Cornell University,
will present the talk titled: “*Advancing biochemical studies with LLMs and
PLMs*”.
The zoom link and abstract are below. We hope you can join us!
Best,
Marcela
--
Wednesday, June 3rd, 1PM ET
10A PT | 11A MT | 12P CT | 1P ET
Find your local time *here*
<https://www.timeanddate.com/worldclock/fixedtime.html?msg=AgBioData+June+20…>
.
Join Zoom Meeting
https://us06web.zoom.us/j/82038356125?pwd=YVFMRElMdEpHZmtObXFvZlA4QVFXQT09
Meeting ID: 820 3835 6125
Passcode: 160683
--
Speaker: Gaurav Moghe, Associate Professor, Integrative Plant Science @ Cornell
University
Title: Advancing biochemical studies with LLMs and PLMs
Abstract: As the number of sequenced plant genomes continues to accumulate
rapidly, innovation in functional annotation of genes is increasingly
becoming a critical endeavor. However, allelic divergence,
duplication-divergence, promiscuity, and redundancy are major hurdles in
function prediction. In this talk, I will describe a workflow we have
developed for enzyme function prediction that includes LLM-assisted
extraction from literature (FuncFetch), organizing in databases
(FuncZymeDB), functional prediction using phylogeny (FuncPred-OG) and using
protein language models (FuncPred-AI). Together, this workflow enables
rapid prediction of substrate classes utilized by BAHD acyltransferases
(our test enzyme family), has high data provenance, is expandable to other
families, and will complement manual biocuration efforts.
--
Marcela Karey Tello-Ruiz, PhD
AgBioData Program Manager
Phoenix Bioinformatics