Planning

Morning session

  • ⏰ 9h-10h: Introduction to Computo and Quarto
  • ☕ 10h-10h30: Coffee break/Discussion
  • 🧑‍💻 10h30-12h00: Hands-on with a toy example

Afternoon session

  • 📝 13h30-15h00: Follow-up with personal article submission process

Learning objectives

  • 🎯 Understand the benefits of reproducible research
  • 📄 Learn how to create a quarto document
  • 📝 Learn how to include code, data, and narrative text in a quarto document
  • 📤 Learn how to submit a quarto document to Computo
  • 🧭 How to navigate the Computo submission process (optional)

Short introduction to Computo and quarto

Team

Chief Editors Associate Editors

Technical Support Community Management

Julien Chiquet (chief editor)

Stat. learning for life science
Paris-Saclay University, INRAE

Pierre Neuvial (chief editor)

Statistics
CNRS, IMT Toulouse

Mathieu Carrière

Topological data analysis
Inria, Sophia Antipolis

Fra.-Dav. Collin

Research Engineer, IA for Microelectronics
LIRMM, CNRS, Montpellier

Aymeric Stamm

Research Engineer, Statistics
LMJL Nantes, CNRS

Marie-Pierre Étienne

Statistical methods for ecology
ENSAI CREST, Rennes

Nelle Varoquaux

ML and causal inference for genomics
CNRS, Grenoble Alpes

Chloé Azencott

ML for therapeutic research
Mines ParisTech

What is reproducible research?

Fundamentally, it provides three things:

Tools to reproduce the results (that’s like cooking)

A “recipe” to reproduce the results (still like cooking)

A path to understanding the results and the process that led to them (unlike cooking…1)

Pre-Computo era

The pdf era and paper submission.

The reproducibility was not a priority:

  • 🛠️ Tools had to be bought, installed, and maintained
  • 🔒 Data and code were not shared (social engineering)
  • ❓ Even methodology details are often missing

Pre-Computo era (2)

And then in the Machine Learning domain, there was distill.pub [1]

  • 📊 State-of-the-art visualizations
  • 🔄 Paradigm shift in scientific publication: “distillation” of complex ideas
  • 💯% reproducible (just a git clone and a few standard commands)

but…

Pre-Computo era (3)

 

engineering was too complex for the average scientist (a lot of javascript, etc.)

In fact, the distill.pub project was discontinued in 2021 [2]

distill.pub

The Rise of the Pragmatic

distill.pub’s goals were right, but they outpaced themselves in terms of development complexity.

  • 🚀 Computo is a fresh start with a pragmatic approach
  • 🧰 Leverage what the scientific community is already using (Rmarkdown, Jupyter notebooks, etc.)

\(\Rightarrow\) bring the community to the higher standards

The Rise of the Pragmatic

distill.pub’s goals were right, but they outpaced themselves in terms of development complexity.

  • 🚀 Computo is a fresh start with a pragmatic approach
  • 🧰 Leverage what the scientific community is already using (Rmarkdown, Jupyter notebooks, etc.)

The Rise of the Pragmatic

distill.pub’s goals were right, but they outpaced themselves in terms of development complexity.

  • 🚀 Computo is a fresh start with a pragmatic approach
  • 🧰 Leverage what the scientific community is already using (Rmarkdown, Jupyter notebooks, etc.)

Origin of Computo (~ 2020s)

French Statistical Society appoints a “publication” committee (led by Julien then Pierre) to develop a new journal

Assessment

  • 😔 Multiplication of “traditional” journals…
  • 😔 No valorization of “negative” results
  • 😥 No or not enough valorization of source codes and case studies
  • 😱 ↘ of publication quality and time dedicated to each article (on author or reviewer sides) [3]
  • 😱 Issue with scientific reproducibility (analyses, experiments) [49]

Point of view

  • 🔄 Need for renewal regarding scientific research implementation
  • 📈 Need for higher standards regarding result publications

⇝ Emergence of “Computo” idea

Philosophy

Scientific perimeter

Promote contribution in statistics and machine learning that provide insight into which models or methods are more appropriate to address a specific scientific question

Open access

  • 💎 “Diamond” open access (free to publish and free to read, possible to reuse)
  • 🅭 🅯 Content published under CC-BY license (attribution, share, adapt)
  • 📝💬 Reviews and discussions available after acceptance for publication (anonymous reviews)

➡️ In accordance with Budapest Open Access Initiative (BOAI) and Plan S

Reproducible

  • 🔢 Numerical (statistical) reproducibility is a necessary condition
  • 💾 Source code and data should be available, at least partly executed and fully executable

Note on reproducible research [1012]

Why reproduce scientific results?

  • 💪 To strengthen their credibility
  • 🔍 To check for errors (everyone makes errors at some point!!!)
  • 🧱 To build new research upon them (science is incremental)

Issues?

  • 🔄 Reproduce numerical scientific results is often difficult (technology/environment evolution, source code/environment configuration/software partially available or not available)
  • ⏳💸 Waste of time and resources to reproduce existing non-reproducible results

Reproducible research?

  • 👥 For others but also for your future self
  • ✅ Improve result credibility
  • 🚀 Facilitate future research works

Setup

Official launch at the end of 2021