So You Have an Idea. Now what?
Abstract
The plethora of open data that is now available creates opportunities for important and innovative science through data synthesis. Synthesis science combines data from multiple studies to draw conclusions about broader ecological theory and mechanisms or to extend the range of inference. It often results in higher-impact publications and generates science that is especially valuable in practice and policy. When done well, it’s also really fun! In this workshop, we’ll explore the interactive process of developing a synthesis project idea and building a team that can see it through. In the first portion of the workshop, we’ll cover common team synthesis models and funding sources, as well as the skills that are essential for gathering a team, keeping them on-track, and facilitating generative discussion. In the second half of the workshop, we’ll break into small groups, facilitated by experienced synthesizers, to explore your synthesis ideas and help you develop plans to move them forward.
Objectives for learners are to
- gain team science and project management skills that are needed for synthesis and that are immediately useful in a research team setting,
- locate and become familiar with resources to gain the data science skills needed for reproducible synthesis
- focus on developing their own synthesis projects, whatever stage they are in, through interaction with instructors and peers.
Code of Conduct
Briefly: Treat all participants with kindness, respect, and consideration, valuing a diversity of views and opinions. Report unacceptable behavior at (844)-641-4133 or email codeofconduct@esa.org
Agenda for 2026-07-27
This agenda is subject to change!
| Timing (PT) | Content |
|---|---|
| 1:30 pm | Welcome & Introductions |
| 1:40 pm | Synthesis Process and Resources |
| 1:50 pm | What makes a “good” question |
| 2:10 pm | Building and stewarding a team |
| 2:20 pm | Locating and managing data |
| 2:30 pm | Discussion |
Introductions
Before we begin, we’d love to get a sense for who you all are and why you’re interested in synthesis work! Please briefly introduce yourself at your tables:
- Your name and pronouns
- A 1-sentence summary of your work
- Briefly, why are you interested in synthesis?
Decide on a name for your table-group and add a tab to the notes document with that name.
Note on Course Materials
This workshop is an adaptation of a full short course that was offered in 2024 and 2025. Materials from those courses are available on the LTER-owned “eco-data-synth-primer” GitHub repository, which also links to the materials for this workshop. So, we recommend that you save the link to this site so that you can have easy access to these materials now and as they are refined going forward. Also, much of the content is covered more completely in our semester-long synthesis course: Synthesis Skills for Early Career Researchers (SSECR).
If you are a GitHub aficionado, we have deployed this website via GitHub Pages so you could also “star” the website’s repository. Simply click the GitHub octocat logo on the right side of the navbar (at the top of the screen) to be redirected to the GitHub repository underpinning this website.
What Do We Mean by Synthesis?
Synthesis: Bringing together data from many sources to generate new scientific insights or test new hypotheses.
Today, we’ll focus on the particular flavor of synthesis that involves a team of people.
Process Overview
Typically, a group of researchers–or researchers and managers or community members–will plan a series of meetings over 2-3 years. The mix of in-person v. virtual meetings and work will vary across different groups and different funders, but the general pattern is similar.

That all seems straightforward enough, but the way it really works is a little more…dynamic. Projects spin off into new territory, promising leads turn out to be dead-ends, key people get new jobs and have to leave the group. Much of what we’ll discuss today will help future-proof your process in the event of changing priorities and changing membership.

From: Ten simple rules for ecological synthesis research (2026) Methods in Ecology and Evolution, in press
Why Synthesize?
- Higher impact products
- Incorporating (and learning) new skills
- Building your science community
- Leveraging the effort of the science community by re-using data
- Inclusion — a way to involve less research-intensive institutions and individuals who can’t/don’t want to do fieldwork
| Academic | Application |
|---|---|
![]() |
![]() |
| What benefits were most important to NCEAS working group participants? | Motivations for synthesis Hackett, E.J. (2020). Sustainability, reproduced from Reid et al. (2009) PNAS |
Funding and Resources
What do you need to “do” synthesis?
- A way to get people together. Unless you decide to work only with people you already know, it’s really helpful to have at least one in-person meeting where you wrestle with objectives and approaches.
- Facilitation skills. Many places to acquire these, but also consider hiring a pro the first time around.
- Project management skills. Maintaining the interest of volunteer researchers means making very efficient use of their time.
- Data: your own and others’. The most impactful syntheses access a wide variety of data sources that haven’t been previously combined.
- Harmonizing diverse data requires a bit of analytical muscle (your own or others’)
Sources of Support for Synthesis
What organizations do you know that fund synthesis work?
- New Phytologist Workshops
- Gordon Research Conferences
- Chapman Conferences
- British Ecological Society (one-third of participants from the developing world)
- ULTRA-Data DCL
- Core programs, e.g., DEB, IOS, MCB, Emerging Frontiers, OCE, OPP, RISE
- Search solicitations for “synthesis activities,” “synthesis projects”
- NSF workshops
| Funder | Travel | PI salary & direct project support | Direct analytical support | Postdoc funding | Facilitation & communication support |
|---|---|---|---|---|---|
| NSF | ✓ | ✓ | ✓ | ||
| ESIIL | ✓ | ✓ | ✓ | ||
| LTER | ✓ | ✓ | |||
| NCEAS/Morpho/Luna | ✓ | ✓ | |||
| Powell Center | ✓ | ✓ |
Everyone offers travel support, but different programs offer different kinds of additional support — remember to budget for the kinds of support that are allowed.
Identifying a Synthesis-Ready Question
Lots of questions are interesting, but not terribly well-suited for a synthesis approach. We’ve learned through experience that there are a few qualities that make some questions a better fit for a) combining data; and b) work by a group. The main qualities that we seek in synthesis projects include:
- Novel and interesting enough to spend 2–3 years deeply involved
- Ripe for synthesis
- Data already exist
- Large geographic area or temporal extent
- Data that aren’t normally collected/analyzed together
- Clearly framed, but flexible
- Results could include several kinds of products: papers of various kinds, symposia, datasets, analytical packages
- Responsive to the funding call!
Keep in mind that early synthesis activities like data discovery and harmonization do not always go to plan. Successful synthesis sometimes requires reframing the research questions or approach as the project proceeds.
For example, if data prove difficult to find or analyze, the group could consider shifting to higher-level research questions and writing papers that:
- update or propose new conceptual models and provide supporting case-studies.
- serve as a “call to action” regarding new research opportunities or gaps in a field’s observational power.
Examples of Recent Synthesis Groups
![]() |
Controls on River Silica Exports (Combined data from more than 1000 watersheds on 6 continents.): Climate, Hydrology, and Nutrients Control the Seasonality of Si Concentrations in Rivers. JGR Biosciences 2024 |
![]() |
Fire and Aridland Streams: Persistent and lagged effects of fire on stream solutes linked to intermittent precipitation in arid lands. Biogeochemistry 2024 |
![]() |
Soil organic matter: multi-scale observations, manipulations & models: Dataset: SoDaH: the SOils DAta Harmonization database, an open-source synthesis of soil data from research networks, version 1.0., EDI 2020 |
![]() |
Global Synthesis of Multi-year Drought Effects (Integrated 100 grassland and shrubland sites across 6 continents): Extreme drought impacts have been underestimated in grasslands and shrublands globally. PNAS 2024. |
Open our communal notes document, head to your tab and take a few minutes to jot down your own notes, then discuss.
- What ideas do you have for synthetic studies?
- What factors make them synthesis-ready?
- How could they be adapted if needed (narrower or broader scope, alternative products)?
Often, early career researchers will be excited about the idea of synthesis but be unsure how to connect with existing or nascent synthesis efforts. Here are a few ideas for how to make yourself available and valuable to synthesis groups.:
- Make it known you want to be involved in synthesis
- Let your advisor know
- Share your enthusiasm
- Skill building:
- Synthesis Skills for Early Career Researchers: SSECR
- Data Carpentries
- ESIIL: innovation summit, hackathons
- Build your community
- Ask questions at meetings
- Initiate conversations
- Start your own!
Building a Team
Research from organizational development (Horowitz and Horowitz 2007, Agrawal and Wooley, 2018) and the science of team science (Cheruvelil et al. 2014) finds that teams combining members form multiple fields, identities, backgrounds, and functional styles are more creative and produce better results.
Diverse Teams Yield:
- More creative and innovative approaches
- Avoiding ‘groupthink’
- Better cognitive elaboration; more complex thinking
- Checked assumptions
- Diversity broadens group scanning ability and consideration of alternative solutions
- Diverse groups achieve better task completion and more efficient use of resources
- Solutions generated by diverse groups have more legitimacy and salience
What are some types of difference you might seek out among group leaders and members?
The composition of the leadership team will affect the success of the project and who you will be able to recruit to the larger group. Look for:
- Different (and complementary) areas of expertise (ecologists, hydrologists, computer scientists, developmental biologists, engineers, etc.)
- Research approach (modelers, empiricists, humanists, community-engaged researchers)
- Complementary professional networks
- Facilitation and project management skills
- Emotional Intelligence
- Study system
In our experience at NCEAS and the LTER Network Office, we’ve found teams of up to 10 to 15 people to be optimal for synthesis work. As individuals, we all have strengths and weaknesses. The beauty of working in teams is that you can invite people who offset your own weaknesses and who bring strengths you don’t have. Often, you’ll have a few core team members who have generated a synthesis idea, but then you’ll want to take a clear-eyed look at what additional skills and qualities to invite. When you do so, be sure to consider:
- Skills, Aptitudes, and Communication Styles
- Look for a mix of empiricists, theorists, and modellers
- Big-picture thinkers, organizers, task-oriented do-ers
- Deep thinkers and risk-takers
- At least some skilled coders
- Career stage
- Senior investigators connect the team to existing literature and fields of study, connect to a broad network of experienced researchers, and have good knowledge of resources, but often have a very limited amount of time to devote to discussion and analyses.
- Junior team members often bring a fresh perspective, familiarity with newer literature, strong coding skills, and time to devote to the project.
- Emotional intelligence
- Research shows (Aggarwal and Woolley, 2018) that the bump in creativity seen in mixed-gender teams is typically due to an increase in emotional intelligence and attention to team dynamics. Include at least a few people with a process orientation and strong people skills.
- Power dynamics
- You won’t be able to anticipate all of the issues related to power dynamics that can arise, but keep them front of mind as you assemble a team.
- Remember that participation in synthesis represents a significant career opportunity.
- Be mindful that such career-building opportunities have not been fairly distributed.
- Be intentional about seeking out people who may not be part of your typical circles
No Diversity Bonus without Trust
Diverse teams only get those positive results when the team adopts a learning mindset and the collaboration process allows each participant to contribute in the way that works for them (van Knippenberg and Hoever 2017), which requires trust, communication, and coordination.
- Get to know each other. Spend time early to talk through various perspectives on the question that may be present in the group
- Establish group norms collaboratively
- Articulate your coordination practices (Where will you keep data? How will you communicate? How often will you meet? What’s your authorship model? Do these work for slow and fast processors? Introverts and extroverts?)
- Pay attention to speaking time
- Ask the “dumb” questions, they often surface unexamined assumptions
- Spend social time together — meals, activities



Data Sources
One of the most frequent stumbling blocks that we have seen groups encounter is finding that data they thought existed, doesn’t. Identifying – and closely examining – data sources early can help you avoid this pitfall. Some sources of data–such as modern remote sensing products, NEON data, and census data–have very clear, explicit ways to access and download them or work with them in the cloud. But the most interesting synthesis questions often involve combining such “big” data with other data sources that may have been collected manually, by a variety of methods and different technicians, over decades.
What kinds of data sources might you consider including in a synthesis project, in addition to your own or others’ field data?
- DataONE
- Environmental Data Initiative (EDI)
- GenBank (NCBI)
- National Ecological Observatory Network (NEON)
- US Geological Survey Data
- Global Biodiversity Information Facility (GBIF)
- NASA Remote Sensing Data
- US Park Service
- FluxNet
- Phenocam network
- iNaturalist
- eBird
- Census data
- Data extracted from papers
- Scraping social media
- Text analysis
- ….
Data Use Principles
There are a few ethical and practical guidelines that will save you a lot of trouble if you can adhere to them from the start of a project.
- Data sources should always be cited
- Keep track of your data sources (sources (including dataset version), permissions, notes, related metadata, what’s included)
- Also track search criteria
- Keep the data that everyone is using in one place (single source of truth)
- Develop a system for organizing data as a group (locations, hierarchy, file and field naming)
- Communicate with data creators whenever possible
- This doesn’t need to be onerous and it can uncover issues and opportunities associated with data sources.
Dear Dr. Smith,
I am working on a synthesis of soil inverbrate diversity across North America and would like to include your dataset titled “xxx” (doi: xxx). We have downloaded the data from yyy repository, but wanted to let you know we are using it and to inquire whether there is any additional context we should be aware of or related datasets we should be sure to include. A short description of the synthesis project follows. Please let me know by xxx date if you have any questions or concerns with our use of this data.
Thank you,
Dear Dr. Smith,
I am working on a synthesis of soil inverbrate diversity across North America and would like to include the dataset behind your paper titled {xxx} (doi:{yyy}), which seems highly relevant. Would it be possible to obtain the data? Ideally, we would access it through through a public repository, such as the Environmental Data Initiative, which offers assistance in curation and submission of datasets. Either way, we would credit you as the data originator and want you to know we are using it. If there is any additional context we should be aware of or related datasets we should be sure to include, please let us know. A short description of the synthesis project follows.
Thank you,
Keeping Track of Data
In the frenzy of data discovery, it is really easy to jump from one “perfect” data source to another, dumping them all into a folder to sort out later. Resist the urge! Note basic information about each data source as you identify it. It will allow you to see gaps and patterns as they emerge and save many hours of sleuthing later.
- URL to the data (and metadata) source
- Sampling location and site (including both coordinates and associated organizations)
- Short Data Description
- Coverage Dates/Frequency
- Filename (as stored on Google Drive or shared file repository)
- URL to Files (cloud drive, website, server; e.g. google drive link)
- Date the data was last accessed / downloaded
- Data Creator/Owner’s Name
- Data Creator/Owner’s Email/contact
- Working group participant who got the data (Name)
- Used in your analysis? (Y/N)
- Any additional notes or decisions about how the data is or will be harmonized and analyzed
Opinions on the reliability, practicality, and ethics of AI use differ substantially. Some see it as a value-neutral tool, like a programming language or a design application. Others have strong moral reservations about its use. Most groups will need to have at least two kinds of conversations about it. Early in your team-forming process, the group should make some broad decisions about how extensively you want to use AI and for what tasks.
A few potential uses to discuss:
- data discovery
- literature review
- literature synthesis
- original coding
- commenting and cleaning code
- recommending statistical analyses
- drafting figures
- original writing
- editing
- proofing
No matter how extensively the group decides to use AI, there is also a very practical concern about reproducibility. Using AI can dramatically speed data discovery and construction of analytical pipelines, but also presents a strong temptation to just “get it working.” Just as you will soon forget why you chose to filter your data in a particular way or chose one analytical technique over another, you will also forget which model and prompts you used, potentially introducing biases that cannot be unwound at a later time.
The LTER Network has developed a template AI use tracker that may serve as a useful starting point for recording how your group uses AI. Note that this should be used after your group has already agreed on how to use AI tools!
What uses of AI seem important to report? Have you already encountered situations where you wish you had kept better track of your AI use?
To provide some minimum attribution and reproducibility for AI use during synthesis projects, have the group document and preserve at least these things:
- What tasks were assigned to AI tools?
- What models were used, including provider (OpenAI, Anthropic, etc.), model name, and version?
- What specific prompts were used?
- Were other pieces context provided to the model (AGENTS.md, skills, knowledge bases, etc.)?
- What data files were generated or modified with AI assistance?
- Describe any human verification and error checking of AI-assisted products.
Make a Plan
Once the synthesis team moves into the operational phase of research, which includes the integration and analysis of data, there are some key activities that must happen:
- Make a plan for technical challenges
- Harmonize and tidy data in a reproducible way
- Analyze data to answer your questions
Once you have analyzed your data, you can interpret the results and create products as you likely would in a non-synthesis project–though the scale and impact of your products are likely to be larger for a synthesis project than for a typical scientific effort. Module 3 of our 2025 workshop has more detail on some of these approaches. Here we highlight some of the most important considerations and practices for a team science approach to the nuts-and-bolts of synthesis research.
1. Make a Plan
For non-synthesis projects, it can feel intuitive to skip this step. Especially if you are working alone or with a small group of close collaborators, the formal process of making and documenting your strategy can feel like a waste of time relative to ‘actually’ doing the work. However, synthesis work uses a lot of data, requires a high degree of technical collaboration, and involves a series of big judgement calls that will need to be revisited. Because of these factors, taking the time to create an actual, specific, written plan (and periodically updating that plan as priorities and questions evolve) is vital to the success of synthesis projects!
See the sub-sections below for some especially critical elements to consider as you draft your plan. Note this is not exhaustive and there will definitely be other considerations that will be reasonable for your team to add to the plan based on your project’s goals and team composition.
How Will You Get (and Stay) Organized?
Project organization is one of those topics that is–maybe–not exciting but is absolutely critical to the success of synthesis working groups. In individual projects conducted over short timescales, an absence of an overarching strategy for organization can be survivable but in a project that will last for several years and will have many actively contributing colleagues this lack quickly becomes an insurmountable hurdle. Also, it is easier to start organized than it is to get organized, so implementing a system at the start of your project and sticking to it will be much easier than trying to re-organize months or years worth of work after the fact.
When Should We Plan?
It is never too early to draft your organization plan! The longer you wait to discuss an organization strategy and decide on a system the worse it will be to go back and retroactively sort the content your group has produced into the structure you agree upon. If you make a plan early on it should be clear where a given file should ‘live’ as it is created which completely removes the labor-intensive reorganization that is inherent to organizing plans made later in a project’s lifecycle.
Broad Considerations
There is no single “best” mode of organizing a project. However, there are some decision points that most groups consider even if not all groups reach the same conclusions at those junctures. Here are some guiding questions that might prove helpful as your group discusses your organizing plan:
- What “modules” or “silos” are likely to be necessary for your project?
- Semi-discrete subcomponents of the project are likely to exist and should (probably) be separated into different folders
- For example, groups almost always have at least the following four major folders: “data”, “notes”, “publications”, and “presentations”
- What level of organization can your group easily maintain for 2-4 years?
- Ideally your chosen structure will require little to no maintenance after it is initially set-up
- How hard will it be to onboard new members to navigate your chosen system?
- Your group will likely need to onboard new members and if they don’t know where they should add their contributions they may add files in incorrect places or refrain from contributing at all–either outcome would be a sad loss for your group
- Note this question also applies to ‘future you’ if you focus on other work for a time and then need to remind yourself how to navigate this project
LTER Recommendation Example
While there is no “one size fits all” solution, our team has identified a structure that has worked quite well for past groups because it is relatively simple to maintain and easily extensible as project questions evolve. In addition, it avoids an overly nested folder structure which makes it easier for new members (or ‘future you’) to become familiar with the overall schema.
Top-level folders are colored blue so that the high-level structure is easier to quickly scan.
Google Drive Structure
Shared Drive
|– data
| |– data-log.csv
| |– raw
| |– tidy
| | |– 01_data-harmonized.csv
| | |– 02_data-wrangled.csv
| | └ - 03_data-filtered.csv
| |– climate
| └ - land-cover
|– notes
|– presentations
└ - publications
|– community-composition
| |– data
| └ - graphs
└ - synchrony
GitHub Structure
GitHub Repository
|– README.md
|– .gitignore
|– scripts
| |– 00_download-data.R
| |– 01_harmonize.R
| |– 02_quality-control.R
| |– 03_filter.R
| |– 04_stats.R
| └ - 05_graph.R
|– tools
| |– README.md
| |– fxn_calc-beta-diversity.R
| └ - fxn_bookmark-graph.R
└ - explore
|– README.md
|– downs-stats.py
└ - lyon-graphs.R
Check out the tabs below for highlights of these structures!
- Limited use of sub-folders
- Consistent folder/file naming conventions
- Good names should be be both human and machine-readable
- Avoid spaces and special characters
- Consistent use of delimeters (e.g., “-”, “_“, etc.)
- Shared file prefix (a.k.a. “slug”) connecting code files with files they create
- Allows for easy tracing of errors because the file with issues has an explicit tie to the script that likely introduced that error
- Contains inputs to code and outputs from code but not code itself
- Trust that GitHub does a better job of tracking code files than Drive can
- Also remember that duplicating code files and storing them in multiple places is a recipe for heartache as there is no “single source of truth” upon which to depend
- Dedicated place for notes / presentations
- Shared
data/folder agnostic to specific product- Will let your group start with the same data product for each paper (just filtered/analyzed differently for each product)
- Contains code but not inputs/outputs
- Use the
.gitignorefile to regulate what Git will/won’t track (for more info, see here)
- Use the
- Dedicated “README” files containing high-level information about each folder
- Numbered script names making workflow order explicit
- Use of custom functions for repeated operations
- This is a great way of ensuring reproducibility–and if you create enough functions, a software package might be a nice ‘bonus’ product for your group!
explore/folder for ‘rough draft’ scripts developed by particular members- Including this can be a great way of making everyone feel more comfortable contributing–even if they are not completely confident in their coding skills
- Including surnames in file names here can be a nice way of avoiding merge conflicts
- You can always rename these files and move them to the
scripts/folder if they seem valuable for the core workflow!
How Will You Collaborate?
Once you’ve decided on your organization method, you’ll need to decide as a team how you will work together. It is critical that your whole team agrees to whatever method you come up with because it will result in a lot of unnecessary work if a subset of people do not follow the plan–and thus introduce disorganization and inconsistency that someone will have to spend time fixing later.
A good rule of thumb is that you should plan for “future you.” What collaboration methods will you in 6+ months thank ‘past you’ for implementing? It may also be helpful to consider the negative side of that question: what shortcuts could you take now that ‘future you’ will be unhappy with?
A huge part of deciding how you will collaborate is deciding how you will communicate as a team. When you work asynchronously, how will you tell others that you are working on a particular code or document file? This does not have to be high tech–an email or Slack message can suffice–but if you don’t communicate about the minutiae, you risk duplicating effort or putting in conflicting work.
Where Will Code Live?
Unlike the other planning elements, we (the instructors) feel there is a single correct answer to this: your code should live in a version control system. Broadly, “version control” systems track iterative changes to files. In addition to the specific, line-by-line changes to project files, contributions of each team member are tracked, and there are straightforward systems for ‘rolling back’ files to an earlier state if needed.
As the comic to the right shows–and as all scientists know from experience–you will have several (likely many) drafts of a given product before the finalized version. With a version control system, all the revisions in each draft are saved without needing to manually ‘re-save’ the file with a date stamp or version number in the filename. Version control systems provide a framework for preserving these changes without cluttering your computer with all of the files that precede the final version.
There are several version control options. Our tutorial from 2025 (which we won’t dive into today) focuses on Git as the version control software and GitHub as the online platform for storing Git repositories in a shareable way. Alternatives to GitHub include, e.g., GitLab and GitKraken, etc.), but they all operate in a reasonably similar way.





