The Rise of Massively Collaborative Science: A Deep Dive into AI and Interdisciplinary ‘Mega Author’ Projects in 2026
Published:
If you browse arXiv regularly, you will be struck by the AI papers whose author lists are so long you have to scroll through multiple pages.
Author: Koutian Wu; GitHub: ktwu01
Frankly speaking, in the past we tended to think of research as a form of closed-door toil, where an old professor hunkers down in the lab with a few PhD students. Over the past couple of days I reviewed the torrent of top-conference papers that have been exploding recently, and I realized the way science works changed long ago.
In fact, projects called “mega author” efforts are becoming the absolute mainstream of interdisciplinary research in 2026.
Here is the story.
Back in the day, when you looked at high-energy physics papers, say CERN publishing a paper on the God particle, thousands of authors would follow in its wake. Everyone treated that as a special case of Big Science. After all, they had to build tens of kilometers of particle accelerators.
Think about it: today’s AI field is perfectly reproducing this Big Science model.
Whether it is OpenAI or the top open-source organizations, they hold the compute and they hold the Transformer architecture. What they lack most is the hard-core data that can stump the strongest AI.
The public text crawlable from the internet was devoured long ago. The bottleneck now sits on the tacit knowledge that only real domain experts possess.
To get that data, the AI community invented a new paradigm: “massively crowdsourced evaluation.”
The Rise of Massively Collaborative Science: A Deep Investigation of AI and Interdisciplinary “Mega Author” Projects in 2026
1. Introduction: A Paradigm Shift from “Lone Genius” to “Thousand-Author Papers”
In the annals of 21st-century science, the mid-2020s will be recorded as a watershed. The pattern of scientific discovery is undergoing its most profound structural change since the era of Big Science. In the past, the pages of top journals such as Nature and Science were long monopolized by small elite labs, but catalyzed by the explosion of large AI models, an entirely new form of research, Massively Collaborative Science (MCS), is reshaping the power map of academia.
This report responds to a cultural phenomenon increasingly popular in tech communities, namely “I published a paper in Nature, but I only made a tiny contribution.” This phenomenon is no accident; it is the inevitable product of the AI evaluation crisis. As the capabilities of large language models (LLMs) approach and even surpass the human average, the traditional automated test benchmarks have comprehensively failed. To evaluate machines that are “smarter” than humans, the scientific community has had to return to human intelligence itself, not just programming ability, but deep domain expertise, linguistic diversity, and complex ethical reasoning.
As of early 2026, this trend has produced several landmark “hyper-authorship” projects. The most typical is Humanity’s Last Exam (HLE), which gathered thousands of experts globally to build a “final human exam” that machines struggle to crack, and eventually reached the highest academic perches. For domain experts outside computer science, for linguists, and even for ordinary members of the public with special knowledge reserves, this is not only a window into participating at the scientific frontier, but a “low-code” or even “no-code” shortcut to top-tier academic publication.
This report will thoroughly survey and analyze similar projects active in 2026, covering AI benchmarks, multilingual NLP, AI safety red teaming, and citizen science projects in biomedicine and astronomy. We will dissect these projects’ operating mechanisms, contribution thresholds, and how to obtain formal academic authorship through “tiny contributions.”
2. The Evaluation Crisis: Why AI Needs Human “Tiny Contributions”
To understand why top AI labs so urgently need non-technical participation, one must first understand the “evaluation collapse” crisis currently facing AI.
2.1 MMLU Saturation and “Inflated Scores from Training Data Contamination”
Before 2023, MMLU (Massive Multitask Language Understanding) was regarded as the gold standard for measuring large-model intelligence. However, by 2025, with the arrival of GPT-5-class models, mainstream models scored above 90 percent on MMLU. This does not mean the models truly reached human-expert omniscience; rather, more often the models “memorized” the question bank from the internet during training. This phenomenon is called “data contamination.”
2.2 The Birth of the “Google-Proof” Standard
To break this impasse, research organizations led by the Center for AI Safety (CAIS) and Scale AI proposed a new standard: if the answer to a question can be found in the first three results of a simple Google search (information retrieval), then it is not a qualified test question.
A true intelligence test must be “Google-proof.” For example, rather than asking “In what year did Napoleon lose at Waterloo?”, ask instead: “Considering the mathematical relationship between the rainfall records of Belgium in June 1815 and the marching speed of French artillery at the time, how many hours was the French artillery deployment delayed by muddy ground, and how did that in turn change the window for Marshal Ney’s cavalry charge?”
Questions like this cannot be answered by crawling ready-made answers from the web; they require the integrated reasoning of a historian. Machines cannot generate this kind of data, because machines are themselves the object being tested. Therefore, the only source of data is humans, especially those who hold knowledge that is not publicly available online, such as out-of-print books, oral history, clinical experience, and dialect slang. This is the theoretical basis by which “tiny contributions” lead to Nature.
3. Core Case Studies: Humanity’s Last Exam (HLE) and HLE-Rolling
As the central reference for this query, Humanity’s Last Exam (HLE) represents the highest standard in this field. Although its first phase ended in April 2025 with a landmark paper, the project did not end; it evolved into the more resilient HLE-Rolling version.
3.1 HLE’s Historical Positioning and Methodology
HLE is not just an exam; it is a global intellectual crowdsourcing experiment. Its core goal is to build a closed-set, multimodal, and extremely difficult academic benchmark.
- Scale: 2,500 expert-level questions.
- Contributors: more than 1,000 professors and researchers from over 500 top universities and research institutions worldwide.
- Difficulty: even a skilled human with internet access needs considerable time to solve these, not merely to search them.
- Outcome: the paper is published in Nature or an equivalent top journal, and all question designers whose contributions are adopted are listed as authors.
3.2 The New Opportunity in 2026: HLE-Rolling (Rolling Update Mechanism)
As AI models iterate rapidly, the static HLE dataset faces the risk of being “overfit” by models. Therefore, in October 2025 the organizers launched HLE-Rolling, a dynamic, continuously updated fork.
3.2.1 Participation Mechanism and Threshold
- Non-technical threshold: participants need not know programming or know how to train models. What you need is domain knowledge. If you are a paleontology PhD student, a senior tax lawyer, or a scholar of Sumerian cuneiform, you are the exact contributor HLE most craves.
- Submission channels:
- Email: send proposals directly to
agibenchmark@safe.ai. - Official dashboard: submit via the Dashboard at
lastexam.aioragi.safe.ai.
- Email: send proposals directly to
- Question design requirements:
- Closed-ended: must be multiple choice or short answer (for easy automated grading).
- Unique: the answer must be objective and unambiguous.
- Hidden: the answer cannot sit directly in the text of a public webpage; it must be reached through reasoning.
- Multimodal: image information such as charts, chemical formulas, and microscope slides is encouraged.
3.2.2 Authorship and Academic Returns
Per HLE’s established conventions and the HLE-Rolling documentation, the organizers commit to “work toward co-authorship for subsequent contributors.” This means that if your question is adopted during the rolling update and enters the core dataset of the next version, you will very likely appear in the author list of a future System Report or journal paper. For people outside academia, this is a heavyweight endorsement; for a current PhD student, it may mean a co-author slot on a top-conference paper.
| Feature | HLE (v1) | HLE-Rolling (current) |
|---|---|---|
| Status | Completed (2025.04) | Active |
| Primary goal | Establish a benchmark | Prevent benchmark saturation/overfitting |
| Contribution method | Centralized solicitation | Ongoing solicitation (email/Dashboard) |
| Return | Nature/arXiv authorship | Authorship/acknowledgment in future versions |
4. Guardians of Language: Multilingual NLP and Decolonizing Science
If HLE is the pinnacle challenge of intellect, then multilingual NLP projects maximize breadth. This is currently the most friendly field for non-technical (non-STEM) contributors. AI faces a serious problem of “English-centrism.” To make models perform well in Swahili, Quechua, or Cantonese, the labs urgently need native speakers’ help.
4.1 Masakhane: The Grassroots Miracle of African NLP
Masakhane (meaning “we build together”) is one of the world’s most successful distributed AI research organizations. It breaks the old model of “the West researches, Africa provides data,” establishing a new paradigm of participatory research.
- Nature: a decentralized grassroots research community.
- Ideal for: native speakers of African languages, linguistics students, translators, social activists.
- Core philosophy: participation is contribution. In Masakhane, translating data, organizing vocabularies, and even explaining cultural context at community meetings are all regarded as “intellectual contributions.”
- Authorship culture: Masakhane’s papers (often published at top venues such as ACL, ICLR, EMNLP) are famous for their extremely long author lists. They explicitly oppose “helicopter science” and insist that every data contributor be named on the paper.
- Active projects in 2026:
- Decolonise Science: translating scientific terms (such as “quantum,” “vaccine”) into African indigenous languages. This requires deep linguistic ability, not coding skill.
- MasakhaNER: building named-entity recognition datasets. All you need to do is read the text and mark people, places, and organizations.
4.2 AmericasNLP: The Revival of Indigenous Languages of the Americas
Similar to Masakhane, AmericasNLP focuses on Indigenous languages of the Americas (such as Quechua, Guarani, Bribri, Nahuatl).
- Shared Tasks: this is the “Olympics” of the NLP field. AmericasNLP holds a competition every year.
- Opportunities in 2026:
- Task 1 (machine translation): although building models requires technical skill, building the evaluation set requires native speakers. You can contribute first-hand translated text, or serve as a human judge to assess model output quality.
- Task 3 (translation metrics): a task newly added in 2026, focused on developing evaluation metrics suited to Indigenous languages. This demands deep understanding of linguistic structure rather than pure algorithmic knowledge.
- Return: participants are usually invited to write a “System Description Paper” or serve as co-authors of a “Findings Paper.”
4.3 Mozilla Common Voice and TidyVoice 2026
Mozilla’s Common Voice is the largest open-source speech dataset. Although ordinary contributors (those merely recording audio) are usually only mentioned in acknowledgments, community leaders and language validators have a chance to enter the inner circle.
- TidyVoice 2026 Challenge: a challenge at Interspeech 2026 aimed at cross-lingual speaker verification.
- Contribution point: the challenge builds on Common Voice data and requires substantial cross-lingual metadata cleaning and validation. For users fluent in multiple languages, helping organize and validate this metadata is a potential path to paper authorship.
5. Safety Red Teaming and Bias Bounties: Building Through Breaking
If you are good at “nitpicking,” or have a sharp intuition for social fairness and ethics, then AI safety is your home turf. Here you need not build models; you only need to break them.
5.1 Humane Intelligence and “Bias Bounties”
Founded by Dr. Rumman Chowdhury, Humane Intelligence is a pioneer of nonprofit AI evaluation. They invented the concept of the “bias bounty,” modeled on the hacker world’s “bug bounty” but targeting algorithmic bias rather than code vulnerabilities.
- The big move in 2026: the Zindi pilot
- In the third quarter (Q3) of 2026, Humane Intelligence will migrate its platform to Zindi (Africa’s largest data science competition platform), launching a large-scale pilot.
- Ideal for: sociology students, advocates for minority rights, and domain experts in specific fields (such as agriculture, healthcare).
- Specific tasks:
- Climate adaptation and Indigenous knowledge: test whether AI’s agricultural advice ignores local traditional ecological knowledge.
- Urban-rural map accuracy: test whether AI shows systematic bias when recognizing satellite images of rural areas.
- No-code participation: you do not need to write scripts to attack the model. Typically the platform provides a chat interface or visual tool; you only need to enter a prompt, record the AI’s erroneous answers, and submit a report.
- Return: in addition to cash bounties, winners and submitters of high-value reports are usually invited to co-author “retrospective papers,” which carry high citation rates at AI ethics conferences such as FAccT and AIES.
5.2 MLCommons: AI Risk & Reliability (AIRR)
MLCommons is the “ISO standards organization” of the AI world. Its AIRR working group is dedicated to setting industry standards for AI safety.
- SafeBench competition: a $250,000 prize for new safety benchmark ideas.
- Contribution opportunity: what they need is not just data, but test design ideas.
- Example: if you are a psychologist, you can design a scheme for testing whether AI can manipulate users through “gaslighting.”
- Example: if you are a legal expert, you can design a scheme for testing whether AI would provide advice that violates GDPR.
- Authorship: the white papers and benchmark papers published by the AIRR working group (such as AILuminate) usually include all active working-group members. Joining the group, attending meetings regularly, and contributing ideas is an effective way to establish expert standing in this field.
6. Cross-Disciplinary Science: From Astronomy to Cell Biology with “Citizen Scientists”
Although the user’s query focuses on AI, any mention of “Nature papers” and “large multi-author projects” calls for recognizing the pioneers of citizen science. These fields are the most mature in handling authorship, and their integration with AI technology is growing tighter.
6.1 Zooniverse: The Aircraft Carrier of Crowdsourced Science
Zooniverse is the world’s largest citizen science platform. Projects here typically involve processing massive image or audio datasets, too large for scientists to handle alone, and too imprecise for AI to process reliably yet.
- Active projects in 2026:
- Planetary Response Network (Mozambique floods 2026): marking damaged buildings on satellite maps. This is a classic combination of humanitarian relief and AI training data.
- Snapshot Wisconsin: classifying wildlife captured by infrared cameras.
- How to get authorship?
- Merely doing classification tasks (clicking on images) usually earns only “collective acknowledgment.”
- The secret to authorship: be active in “Talk” (the discussion forum). Scientific discovery often stems from outliers. If you find a strange image (like the famous “Tabby’s Star” or “Hanny’s Voorwerp”) and spark scientists’ attention in the forum, you are likely to be listed as one of the discoverers in the paper’s author list.
- AI error correction: many new projects operate in a “human-AI collaboration” mode. AI does the preprocessing; humans correct its mistakes. This “correction data” is vital for improving model performance, and deep participants (super-users) are often treated as research partners.
6.2 Human Cell Atlas (HCA)
This is an ambitious biology project to map every cell in the human body.
- Contribution method: HCA frequently organizes “annotathons,” inviting medical students and biologists to help annotate cell types.
- Return: HCA is a consortium. Core contributors are absorbed into the consortium, and its papers in Nature and Science typically list a huge “HCA Consortium” author roster that includes data contributors.
7. Deep Analysis: A Return-on-Investment Comparison Across Project Types
To help readers choose the path that suits them best, we compare the above projects across multiple dimensions.
| Project | Domain | Core need (your contribution) | Technical threshold | Authorship probability | Top-journal potential | Ideal audience |
|---|---|---|---|---|---|---|
| HLE-Rolling | AI benchmark | Hard, counterintuitive expert questions | Low (needs domain knowledge) | High (cumulative system) | Very high (Nature-class) | PhD students, industry experts, history/science enthusiasts |
| Masakhane | NLP | African language translation, corpus collection | Low (needs language ability) | Very high (community culture) | High (ACL/ICLR) | Linguists, multilingual speakers |
| SemEval 2026 | NLP evaluation | Humor judgment, text relevance annotation | Low (needs reading comprehension) | Medium (Shared Task paper) | Medium-high (ACL Workshop) | Humanities students, creative writers |
| Humane Intel. | AI safety | Nitpicking, finding bias, attacking models | Low (needs critical thinking) | Medium (retrospective report) | Medium (FAccT/AIES) | Social science background, rights advocates |
| Zooniverse | Natural science | Finding anomalies, correcting AI classification | Very low (needs patience) | Low (must become a super-user) | High (astronomy/biology top journals) | Enthusiasts, students |
| FrontierMath | Mathematics | Proposing unsolved/extremely hard math problems | Very high (needs mathematical depth) | High | Very high (math/AI top journals) | Math graduate students |
8. Practical Guide: How to Maximize the Visibility of Your Contribution
“Making a tiny contribution” does not automatically equal “appearing on the author list.” In large collaborative projects, you need to manage your participation strategically. Here is a survival guide based on 2026 community norms:
8.1 Seek “Meta-Tasks”
Do not only do the most basic data production. Most people only “answer the questions.” If you can be the one who “sets the questions” or “grades the papers,” your standing rises sharply.
- Documentation writing: volunteer to write the Data Sheet or Model Card for a dataset. This is dreary work highly valued in the AI ethics field but that engineers are often reluctant to do.
- Quality control: in Masakhane or Common Voice, take the initiative to serve as a “validator,” reviewing others’ submissions.
- Community coordination: help organize weekly meetings, write meeting notes, manage the Discord channel. This kind of “invisible labor” is highly respected in decentralized organizations.
8.2 Seize the “Opt-in” Moment
Large projects (like BigScience) usually distribute an Authorship Opt-in Form before publishing a paper.
- Beware: this form is typically posted only in the core Slack/Discord channels, or sent via the mailing list. If you quietly submit data without checking the group messages, you will not only miss authorship, you may even miss the acknowledgment.
- Strategy: stay active in official communication channels, and make sure organizers know that behind your ID is a real person.
8.3 Aim for “Shared Task” Paper Opportunities
In competitions like AmericasNLP or SemEval, as long as you submit valid results (even at baseline level), you are usually qualified to submit a System Description Paper. This paper gets included in the workshop proceedings of top venues such as ACL and is a formal academic publication indexed by Google Scholar. This is the most reliable path to “publishing a paper.”
8.4 Leverage the Advantage of “Long-Tail Knowledge”
Do not crowd into “general knowledge.” HLE does not need more questions about “photosynthesis”; ChatGPT has that memorized cold.
- Differentiated competition: contribute knowledge that is not in books and not searchable on the internet.
- Example: specific weather expressions in your hometown’s dialect slang.
- Example: the troubleshooting procedures in the maintenance manual for the specialized equipment at your factory.
- Example: ad copy from the Republic of China-era newspapers in your collection.
- This “private data” or “tacit knowledge” is exactly the nourishment AI currently lacks most.
9. Conclusion: In This Era, Everyone Is AI’s Teacher
Although the user’s query carried a joking tone, it keenly captured a shift in the scientific paradigm. AI is undergoing a process of moving from “data mining” to “data farming.” In this process, machines no longer need more GPUs; they need higher-quality human feedback (RLHF) and human evaluation.
Projects like Humanity’s Last Exam (HLE) are in effect building a digital breakwater for all human knowledge. They prove that the depth, complexity, and unpredictability of human intelligence remain out of machines’ reach. By participating in these projects, whether submitting a carefully designed history riddle or recording a greeting in an endangered language, you are not only contributing to a Nature paper, but also helping define for AI what is “real,” what is “correct,” and what is “human.”
In 2026, the gates of science are open as never before. You do not need to be a PhD, you do not need to be a programmer; you only need to bring your curiosity, your expertise, even your biases (as test samples), and join Discord, Slack, or Zooniverse. In that future Nature paper’s author list, perhaps your name really will appear.
Appendix: Core Resource Navigation
- HLE-Rolling submission:
lastexam.ai/ email:agibenchmark@safe.ai - Masakhane community:
masakhane.io(joining its Discord is strongly recommended) - Zindi competition platform:
zindi.africa(watch for the 2026 Q3 Bias Bounty) - AmericasNLP:
turing.iimas.unam.mx/americasnlp/ - SemEval 2026:
semeval.github.io/SemEval2026/ - Zooniverse:
zooniverse.org - Scale AI Human Frontier Collective:
scale.com/careers(search for HFC Fellow)
Continuing from the above. It is not only a carnival for humanities and computer science scholars; for friends working on climate and geoscience, this is absolutely a blue ocean.
I used to think the spatiotemporal scale of Earth science was too large. To validate carbon-flux models of the Amazon rainforest or Siberian permafrost, you depend on field sampling all over the world. But now this inherent physical distribution has become a huge advantage.
For example, there is the Project Polyclimate project that is the talk of the town right now. Inspired by crowdsourced problem-solving in mathematics, they are GitHub-izing climate science directly. Want to know the world’s remaining carbon budget? You no longer have to wait for lagging government reports. Everyone submits Pull Requests directly on GitHub to fix data sources or optimize computation code. As long as your code is merged, you are a core author of the project’s white paper.
This “code-as-paper” model is absolutely thrilling.
Not only that, many traditional data-collection projects are waking up too. Take Snapshot USA, which runs camera traps nationwide, or the EXCHANGE Consortium, which does water-quality sampling. In the past, if you gave them data, they might at most mention you in acknowledgments. Now it is different: as long as you provide high-quality exclusive data, you can be listed legitimately as a Consortium Author on a top journal like Ecology.
Needless to say, there are hub organizations like Climate Change AI (CCAI). Not only do they run workshops at top venues like NeurIPS and ICLR, they also directly issue seed grants of several hundred thousand dollars. Take a look around their Discord channel and you will find posts everywhere like “I have high-resolution satellite methane data, urgently seeking a computer-vision collaborator.”
Even KDD, the top venue for data mining, opened a dedicated AI for Sciences track in 2026. That is almost as close to writing “we need interdisciplinary experts” on their faces.
Whether you study weather simulation or soil cycles, as long as you can bring your exclusive data and tacit knowledge to feed these large models, they will welcome you with open arms. This is a win-win. They get high-quality evaluation data; you get a shared first-author or core-author seat on a top-conference paper.
For a moment I was left speechless. The advance of technology always sweeps everyone in in ways we never anticipated.
When we are all anxious about whether AI will replace us, AI is desperately calling out for humanity’s last expert knowledge. Machines can generate ten thousand seemingly fluent survey articles in a second. But they cannot create a question that only a geologist knows where the pitfalls are.
No matter how large models evolve, the truth of scientific research still lies in the mass collaboration of human curiosity.
If you are also an earth-science researcher who has slaved away in your field for years, why not go wander around those AI communities today. Take the tacit knowledge in your head and feed this era’s AI leviathans. Smooth out some information gaps, and who knows, you too might leave your name on that long mega-author list.
