Open science
Open science is a concept that aims to make scientific research and data available to both the scientific community and non-expert audiences. It has multiple benefits, such as increasing stakeholders’ engagement, enhancing scientific collaboration, and ensuring the reproducibility of the work along the way. As open science practices gain traction across various fields, this deliverable outlines the development and application of an open science protocol to be followed by the different modelling teams within the IAM COMPACT consortium.
IAM COMPACT comprises a wide range of models and modelling teams, making transparency and thorough documentation crucial to ensuring the accessibility, clarity, and reproducibility of research outcomes to various audiences.
Our open science protocol includes the key open science principles, including the widely adopted FAIR principles—Findable, Accessible, Interoperable, and Reusable. While these principles have been successful in promoting data exchange and usage in multiple projects, there are evident limitations to their effectiveness. To address these gaps, the TRUST principles—Transparency, Responsibility, User focus, Sustainability, and Technology—were developed to evaluate the trustworthiness of digital repositories. Together, these frameworks ensure that research outputs can be the basis of future scientific work, enabling broad access and reuse.
The first stage of the protocol focuses on the documentation and harmonisation of models and data. It is essential to ensure that models are understandable, and the modelling assumptions are recorded for future enquiries. To achieve this, we recommend each team to faithfully describe their models in dedicated platforms (such as GitHub pages[1] or customised wiki pages, e.g., IAMC wiki[2]), detailing the inputs, outputs, assumptions, and specificities to model versions. Whenever possible, each model version should be linked to a release with an associated Digital Object Identifier (DOI). This ensures the model version is easily discoverable, accessible, and reusable by the research community, while also enabling proper citation in future studies.
In the context of IAM COMPACT, all model documentation must also be hosted on the I2AM PARIS platform, which was initially developed in the PARIS REINFORCE project (coordinated by NTUA). Based on its initial design, the platform “seeks to enable modellers to communicate with one another and stakeholders to interact with modelling assumptions, scenarios and results in an informative way and to understand which decarbonisation pathways are the most relevant and realistic, ultimately enhancing the legitimacy of the scientific processes and improving the transparency of the employed methods, models and tools”. This interactive and user-friendly documentation hub will enable other consortium members and stakeholders to better understand how the models work and how they can be utilised, ensuring consistency and clarity across all documentation.
For model interconnections, tailor-made documentation must be uploaded to the appropriate site. On the one hand, the model interconnection, assumptions and input/output interaction can be detailed in specific heatmaps in the I2AM PARIS platform. On the other hand, they should be described in the methodology or Supplementary Information (SI) site of the related publications, which must specify the model versions used. This will promote that models developed by different consortium members can work together in a transparent way.
Finally, to ensure that outputs are comparable models must homogenise their inputs in coordination with D4.3[1] and its updated version D4.4[2], which define a “broad scenario logic”. It describes the common time resolution, input information, or the baseline specifications that modelling teams should implement to the extent possible.
The second stage of the protocol focuses on ensuring the quality and the transparency of produced outcomes. The starting point is the homogenisation of the model results using the time-series data template[3], as described by D4.4 and the nomenclature[4]initiative. This facilitates the cross-comparison of the outputs obtained by different models and modelling teams. Moreover, it enables an automated vetting procedure through the automatised I2AM PARIS validation tool[5]. This platform asserts that the model outputs are reliable by comparing early modelled periods to observed data and by ensuring that key balances match, such as demand and supply values or trade flows. Modelling teams can also make use of manual procedures (although not recommended) or model-specific tools (e.g., gcamreport[6]). In parallel, model projections require independent (“within-consortium”) expert assessment. This process should include a review of the model inputs, assumptions, and outputs. The review should be done in a transparent and collaborative way, with feedback provided to the consortium members to improve the models.
To boost the visibility and impact of key findings and analyses, model outputs should be made available under a DOI. This ensures full transparency, accessibility, and reusability of the data, allowing it to serve as a foundation for future analyses. If the raw output data is too large to be uploaded, at least the relevant study-specific results and/or standardised outputs should be deposited in a secure and accessible platform, such as Zenodo. In the same way, analysis and figure development code must be traceable and open, in GitHub, Zenodo, or other open-source platforms. These datasets must be accompanied by comprehensive metadata, detailing the contents of the repository, the data-handling processes, and guidelines for reuse. This approach not only preserves the methods and data, but also facilitates its findability, broader use, and code recycling in the research community.
The communication of results could also include visualisation tools to effectively disseminate the findings to non-expert audiences. Although developing such tools can be time-consuming, they attract interest from policymakers and stakeholders and ensure full transparency and acceptance of the results. Custom websites or R Shiny applications[1] are excellent resources for creating these interactive visualisations. Finally, we note that, in order to ensure the full accessibility and reproducibility of the results, we are exploring the potential use of additional software platforms (e.g., Docker, Jupyter Kubernetes, etc.) to solve potential problems related to the reproducibility of the execution environment of the models (when licences allow it) and of the model intercomparison and post-analysis studies.
To guide all open science practices of the project, IAM COMPACT has produced two reports detailing a protocol for open science:
Rodés-Bachs, C., Sampedro, J., Horowitz, R., & Van de Ven, D.-J. (2023). IAM_COMPACT_D3.6_Open_Science_Protocols. Zenodo. https://doi.org/10.5281/zenodo.10464845
Rodés-Bachs, C., Horowitz, R., & Van de Ven, D.-J. (2024). IAM_COMPACT_D3.7_Open_Science_Protocols_update. Zenodo. https://doi.org/10.5281/zenodo.13957129
A presentation detailing IAM COMPACT’s Open Science Principles can be found here. The aim of this presentation was to disseminate among modelling partners the updated Open Science Principles protocol as well as demonstrate valuable open-source platforms and tools relevant to various protocol steps, among them a project-specific tool for vetting and validation of modelling outputs (available here).
[2] https://www.iamcdocumentation.eu/IAMC_wiki
[3] https://doi.org/10.5281/zenodo.10430891
[4] https://doi.org/10.5281/zenodo.13132311
[6] Daniel Huppmann, Laura Wienpahl, Philip Hackstock, & Lucie Castella. (2023). Nomenclature—Working with IAMC-format project definitions (v0.9.1). Zenodo. https://doi.org/10.5281/zenodo.7956229