Thanks for a great July Meeting! Stay tuned for what’s next from ESIP.

How to Build a Living Document – the Bio Data Standards Primer Guidelines

How to Build a Living Document – the Bio Data Standards Primer Guidelines

Note: this blog is edited and adapted from “Building a Living Document: The Story,” part of the Primer Guide itself. Visit the full story for more information.

The ESIP Biological Data Standards (BDS) Cluster Primer Guide is a community-driven, living resource designed to help navigate the complex ecosystem of biological data standards. Maintained by the BDS Cluster at Earth Science Information Partners (ESIP), this web-based guide serves as an accessible introduction to best practices for data context, integration, interoperability, web-readiness, and software integration. It turns the inherent complexity of biological data from a major operational bottleneck into a driver for collaborative, scalable Earth and ecological science.

There was a lot to consider when building this resource, over many years. And the work is not yet done. Many lessons were learned and, more importantly, documented. If you want to build your own similar resource, look no further than the BDS Primer Guide for an example of a true collaborative effort leading to a practical, accessible product.

An illustration of some of the many types of biological data that are typically created to observe ecosystems on plante Earth.

The Problem: Biological Data Standards Are Hard to Navigate

Biological data are messy by nature. They come from acoustic sensors, visual surveys, genetic sequencing, camera traps, and a myriad of other biological observing methods. Because these sensors capture data in different formats, each project tends to organize its data in its own way. Standards exist to help; but, for data managers new to the space, the landscape of metadata standards, controlled vocabularies, taxonomic authorities, and web services can feel overwhelming.

The ESIP Biological Data Standards Cluster was formed in the wake of the 2020 ESIP-IOOS Biological Data Standards Workshop. The Cluster's mandate included developing guidance, best practices, training, and building a community to support the standardization and sharing of biological data. 

People showed up to the collaboration from ESIP’s federal funding agencies (USGS, NOAA, and NASA), partnering scientific networks (e.g., OBIS, GBIF, NEON), academic institutions (e.g., Florida State University, Alfred Wegener Institute), and independent organizations (e.g., MBARI, Intertidal Agency). The energy was real and palpable; they all cared about the complex problem and wanted to contribute.

This energy was channeled into a product that could evolve over time, rather than becoming a static artifact the moment it was published. It is not a step-by-step tutorial; platforms change too quickly for that. Instead, it is a set of principles established for designing products that any open science community can adapt.

So how did this resource get to the version it is now: a living document? How did the Cluster take the steps and effort to go from a static Primer to an evolving and adaptable Guide?

Act I: The Primer (2020–2022)

The Cluster's first idea was to build a decision tree — a flowchart to guide data managers to the right standards needed for their situation. But, biological data proved too heterogeneous. As the Cluster's founding chair Abby Benson described it, the decision tree quickly became a “decision forest.”

Instead, the Cluster decided to create an infographic-style primer: a visual, one-page-ish overview of the standards landscape, structured around the FAIR data principles, which were relatively new at the time. The Biological Observation Data Standardization: A Primer for Data Managers was published on Figshare in 2021, designed to be printed at conferences and shared as a quick reference. It covers metadata standards, data standards, taxonomic authorities, habitat classifications, web-enabled standards and services, and some baseline best practices.

The Primer was a success. It filled a real gap, generated strong discussion at ESIP meetings, and won recognition. Benson received the 2022 ESIP Catalyst Award in part for this work. However, at the time, it was – by design – a snapshot. There was room for improvement.

As a static PDF with limited details, the resource didn't tell the user much about how to implement the standards it referenced. The Cluster recognized that future iterations required something more than a list of links, but wanted something that was not difficult to maintain.

Act II: The Guidelines & Continuous Improvement (2022–Present)

The Spark

In late 2022, the Cluster identified the next step: a companion product that would expand each section of the Primer into actionable guidance. Where the Primer said “use Darwin Core,” the Guidelines would explain what Darwin Core is, why people use it, and how to get started. Where the Primer listed best practices, the Guidelines would provide context, examples, and references.

Cluster co-chairs made a critical design decision early on: this product would be a living document, built on digital infrastructure that made feedback easy to give, contributions easy to track, and updates easy to publish.

The process of creating the document was as important as the document itself. If the process was open, collaborative, and low-friction, then the product could continuously improve with community input. If the process was closed or cumbersome, the product would stagnate regardless of how good its first version was.

Capturing Contributions

There was broad interest in contributing to the Guide from Cluster participants, who came from different institutions and career stages. The challenge became how best to capture and incorporate these contributions. How do you focus the ideas and energy of busy professionals, most of whom volunteer their time, into a coherent and evolving product?

The co-chairs selected a few useful tools and practices that made it easy for people to contribute in whatever way suited them, while keeping all contributions organized and traceable. The key components were:

  • GitHub as the backbone. The bds-primer-best-practices repository became the single source of truth. All content, all feedback, and all decisions could live there, and GitHub offers version control, issue tracking, pull requests for review, free hosting for the rendered site, and more.
  • GitHub Issues as a low-barrier entry point. Anyone with an account can open an issue to suggest a change, flag an error, or propose new content. Issues were labeled and organized using milestones so that contributors could identify what was being prioritized. 
  • Google Docs as a drafting bridge. Not everyone is familiar with GitHub, so the Cluster used Google Docs as an additional collaborative drafting space. This two-track approach met people where they were without sacrificing the benefits of version control.
  • Slido polls for collective decision-making. When the Cluster needed to make a choice they used live polls. This gave every attendee an equal voice and created a transparent record of how decisions were made.
  • Quarto for rendering, GitHub Pages for hosting. The Guidelines were written in Quarto (a markdown-based publishing system), rendered to HTML via GitHub Actions, and hosted for free on GitHub Pages at esipfed.github.io/bds-primer-guidelines. This means that every merged pull request automatically updates the live site eliminating the need for manual publishing steps.
  • Monthly meetings as the heartbeat. Regular, monthly meetings served as community updates, working sessions to resolve GitHub issues, and time to keep momentum even when individual contributors were busy.
  • Cross-Cluster collaboration. The Cluster didn't work in isolation. For instance, they held a joint “jam session” with the Semantic Harmonization Cluster to bring in expertise on controlled vocabularies and ontologies. Connections to the Marine Data Cluster, Sustainable Data Management Cluster, and Physical Samples Cluster also enriched the content and expanded the contributor base.
  • Teaching GitHub by doing it. At the January 2025 ESIP Meeting, the Cluster hosted “Using GitHub for Collaborative Review of the Biological Data Primer Guidelines,” a session which taught participants to find, tag, and assign issues, plus submit pull requests. They then submitted real feedback, turning it into a contribution event.

Solving the Authorship Problem

Community-authored documents have a persistent social problem: who gets credit? In traditional academic publishing, authorship is a fixed list that may not reflect everyone's contributions fairly. People who are shy about claiming credit can get left off while those who contributed minimally may be listed. And, once the author list is set at publication, there is no mechanism to recognize ongoing contributions.

The Cluster developed an elegant solution: in the GitHub repository, each contributor fills out a YAML template describing their own contributions using the CRediT (Contributor Roles Taxonomy) framework. The template includes clear criteria for what constitutes authorship, so the expectations are transparent and self-service. You don't have to ask to be included; you include yourself by documenting what you did. And if someone claims authorship without meeting the criteria, the record is public and auditable.

A Python script run via GitHub Actions then translates the YAML contributor data into two standard citation formats: a Citation File Format (`.cff`) file and a `.zenodo.json` file for Zenodo deposits. This means that every time the Guidelines are published or deposited, the citation metadata is automatically generated from the contributors' self-reported roles. No manual list-building, no forgotten names.

This system solved three problems at once: it lowered the barrier to claiming credit, it established clear and fair criteria, and it automated the translation to standard formats that archives and citation managers understand.

Intentional Design Choices

Several other design choices are worth highlighting:

  • Aesthetic review. The Cluster invited a USGS product designer to lead a session on visual identity. This reflects a key principle: if you want people to use your product, user experience of encountering it matters.
  • Structured content with a consistent template. Each section of the Guidelines follows a common structure: a preamble explaining the purpose, with key knowledge points, each with a description of what it is, why people use it, and resources.
  • Versioned releases tied to milestones. Rather than treating the Guidelines as perpetually unfinished, the Cluster uses GitHub milestones to define release targets (V1.0, V2.0). This provides a rhythm of “good enough to share” moments while keeping the door open for continuous improvement.
Cartoon included in thh Primer Guidelines (Credit XKCD)
  • The “coalition of the willing” ethos. The work was never mandated by any institution. People contributed because they saw the need and wanted to help. The Cluster's role was to create a structure that made those contributions count in a productive way, not to recruit reluctant participants.

Making your own living document? Here’s what worked

If your community is trying to build a living document collaboratively, here are the principles that made this work possible:

  1. Use a platform that supports both collaboration and publication. GitHub (or a similar platform) provides version control, issue tracking, review workflows, and free hosting in one place. The rendered site is the public face; the repository is the backend.
  2. Meet people where they are. Not everyone will be comfortable with your chosen platform on day one. Provide bridge tools (like Google Docs for drafting) and invest in onboarding (like the Winter 2025 GitHub workshop). The goal is to lower the barrier to contribution, not to gatekeep by technical skill.
  3. Design for feedback, not just publication. If you want a living document, feedback must be as easy to give as possible. A “Give feedback” link on every page, a well-organized issue tracker, and regular meetings to review feedback together all signal that input is genuinely wanted, not just tolerated.
  4. Automate what you can. GitHub Actions can build and deploy your site, generate citation files from contributor data, and run validation checks. Every manual step you eliminate is a step that won't become a bottleneck when the original authors move on.
  5. Make authorship self-service and transparent. Use a framework with clear criteria and let contributors declare their own roles. Automate the translation to standard citation formats. This removes social friction, rewards participation, and produces machine-readable metadata.
  6. Build a rhythm. Monthly meetings, even short ones, maintain momentum and community. Alternate between working sessions (drafting content, resolving issues) and community sessions (invited talks, cross-Cluster collaborations, updates from the field).
  7. Capture the energy, don't manufacture it. The most important ingredient is people who care about the problem. Your job is to design a mechanism that focuses their contributions into something durable. If you don't have that, no amount of process will compensate.

Where It Stands

As of early 2026, the Biological Data Standards Primer Guidelines are published and actively maintained at esipfed.github.io/bds-primer-guidelines. The Cluster continues to meet monthly, with new members joining regularly. The leadership of the Cluster has itself been a living and changing thing over the years, each one equipped by those before them to carry the effort forward. 

Next priorities include responding to the rapidly evolving conversation around AI-readiness of biological data, deepening connections with international standards bodies like TDWG and OBIS, and continuing to refine the content based on community feedback. Specifically, we would like to implement multilingualism and discover ways to streamline contributions from people who are prohibited from using GitHub. We are already working on syncing up development with the Primer, and the infrastructure is in place to do all of this without starting over. The product is not finished… and that is the point.

Acknowledgements

This work was possible because people showed up — month after month, often as volunteers — and gave their expertise to a shared cause. Many people participated in Biological Data Standards Cluster meetings and activities during the period documented here (November 2022 through February 2026), with contributions ranging from leading breakout sessions and drafting content to asking the right question at the right time. Visit the original article for a full Acknowledgements list.

ESIP logo color transparent
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognizing you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.