
From declaration to decision: implementing responsible research assessment
Responsible research assessment has moved well beyond a fringe conversation about the misuse of research metrics.
DORA has existed for more than a decade. CoARA has established a shared direction for reform based primarily on qualitative judgement, supported by the responsible use of quantitative indicators. Funders worldwide are redesigning CVs, changing assessment criteria, limiting the use of publication metrics, and asking reviewers to consider a broader range of research contributions. We know that research assessment is changing. However, the real question is how these changes are translated into practice.
In my previous discussion of responsible research assessment, I argued that changing assessment criteria is not the same as changing assessment practice. Reviewers cannot be expected to infer why a system has changed, and familiar assessment habits can easily survive inside newly designed processes. For responsible research assessment to effectively shape decisions about researchers and research, certain conditions need to be in place.
In 2026, DORA, the Global Research Council and Science Europe released A Practical Guide to Implementing Responsible Research Assessment at Research Funding Organizations. The guide deliberately shifts the conversation from commitment to implementation across the funding lifecycle, including establishing funding programmes, making funding decisions, setting grant conditions, monitoring, and engaging with research communities.
CoARA has also released a 2026 framework on Improving Practices in the Assessment of Research Proposals. Drawing on experiences from different funders, it provides strategic guidance, questions and emerging practices rather than prescribing a single model for reform.
This shift matters because responsible research assessment cannot simply be inserted into a policy document. It must be built into the information applicants provide, the instructions and training reviewers receive, the rules governing assessment, how reviewers interact with one another, and how funders evaluate whether reforms are working.
What implementing responsible research assessment requires
One of the core CoARA commitments is that research assessment should be based primarily on qualitative judgement, for which peer review is central, supported by the responsible use of quantitative indicators. This sounds straightforward, but in practice, it creates a significant design challenge.
If we want reviewers to make more qualitative judgements about research quality, contribution, significance, leadership, or impact, then we need to give them the information they need to make those judgements.
We also need to ensure that:
- applicants understand what evidence they should provide;
- application forms give them enough space to provide it;
- reviewers understand what they are being asked to assess;
- reviewers are trained to distinguish relevant evidence from familiar proxies;
- the amount of information does not create an unreasonable burden on reviewers;
- reviewers are assessing a sufficiently common evidence base;
- rules around inappropriate information are applied consistently;
- opportunities exist to calibrate and challenge individual judgements; and
- implementation is evaluated through assessment behaviour, not simply through the existence of a policy or new form.
These elements are interconnected and weakness in one can undermine the others.
Why implementing responsible research assessment starts with purpose
Reviewer guidance often tells people what to do.
- “Do not use Journal Impact Factors.”
- “Consider research outputs beyond publications.”
- “Assess researchers relative to opportunity.”
- “Focus on research quality rather than publication quantity.”
But reviewers also need to understand why these instructions exist. Without that understanding, reform can appear arbitrary. A new form cannot explain reform on its own.
A reviewer may interpret a Top 10 publications section simply as a shorter publication list. A narrative CV may appear to be a traditional CV rewritten into paragraphs. Restricting the use of Journal Impact Factors can look like an instruction to ignore information that the reviewer still believes is valuable. People will inevitably bring existing mental models into new assessment systems.
If reviewers understand that a reform is intended to prevent journal reputation substituting for assessment of the research itself, the change makes more sense. If they understand that a narrative CV is intended to recognise contributions that conventional publication lists can overlook, they can interpret the information differently.
This is why awareness and training should not be treated as minor implementation activities. CoARA’s qualitative assessor training resource identifies both dedicated reviewer training and assessor-to-assessor exchange before and during assessment as potentially valuable mechanisms for improving qualitative assessment.
Useful research exists in this area. For example, the Research on Research Institute’s (RoRI) work on narrative CVs found that implementation is shaped by the CV format and the environment in which reviewers use it. Its recommendations to funders include aligning criteria with the objectives of narrative CVs, briefing panel chairs, providing guidance to assessors, discussing resistance and considering panel composition.
NHMRC’s 2025–2030 Action Plan for Responsible Research Assessment similarly includes ongoing guidance for applicants and reviewers, evaluation of reviewer support and training, and review of reviewer feedback to determine whether assessment practice aligns with its reform commitments.
Qualitative research assessment needs enough space
There is another implementation issue that I think deserves more attention. If we expect applicants to demonstrate quality through qualitative evidence, we must give them enough space to provide it. The design of an application does not always match the complexity of the judgement it asks for. The NHMRC Investigator Grants publications criterion provides a useful example. In the 2027 round, applicants can nominate up to ten publications. For each publication, they receive a maximum of 1,000 characters to explain why it was nominated, its quality and contribution to science, and their own contribution to the publication.
At the same time, NHMRC asks reviewers to assess publication quality through characteristics including the rigour of experimental design, appropriate use of methods and statistical approaches, reproducibility of results, analytical strength of interpretations and significance of outcomes. Reviewers must also consider the contribution to science and the applicant’s contribution to each publication. NHMRC’s guidance for peer reviewers on assessing publications reinforces this distinction between assessing the research itself and relying on publication or journal metrics.
That is a sophisticated qualitative assessment task, and 1,000 characters is very little space to demonstrate it.
For an original research study, relevant information might include:
- what question was investigated;
- why the question mattered;
- the research design;
- methodological strengths;
- the population, sample or dataset;
- the analytical or statistical approach;
- the principal findings;
- what was novel or significant; and
- what the applicant personally contributed.
I am not suggesting that every applicant should reproduce the methods section of every publication, the issue is proportionality. If an assessment criterion asks reviewers to judge research design, methods, analytical strength, significance and applicant contribution, applicants need sufficient space to identify the evidence that allows those judgements to be made.
Otherwise, compressed statements become attractive:
“Published in a leading journal, highly cited, FWCI 7.8.”
That statement uses very little space, but it tells the reviewer almost nothing about whether the research design was rigorous, the analysis appropriate, the findings robust or the applicant’s role substantial.
The structure of the application can therefore steer applicants back toward the very proxies that responsible research assessment seeks to reduce. Qualitative assessment does not simply require narrative. It requires enough well-structured narrative to make the requested judgement possible.
What current funders are doing differently
No universal amount of narrative space will solve this problem. Different funding organisations have taken markedly different approaches.
These schemes are not directly comparable. But the comparison points to an important design principle. The information requested, the space provided and the judgement expected of reviewers should be considered together.
The Swiss National Science Foundation, for example, asks applicants to describe only one to three major achievements. It discourages lengthy publication lists and instructs applicants not to use citation metrics, journal rankings, institutional rankings or h-index-type proxies. Achievements can cover knowledge generation, development of others, and contributions to wider society.
Research Ireland takes another approach. Its 2026 Stage 2 process allows a five-page narrative CV for the lead applicant and any co-applicant, alongside separate sections for the research programme and impact statement.
UKRI’s R4RI commonly provides 1,150 words distributed across four broad areas of contribution, with a further 500 words available for contextual information rather than additional achievements.
Different systems are testing different answers to the same problem: how do we give reviewers enough meaningful information without simply making applications longer?
Reducing reviewer burden without reducing evidence
One concern about qualitative assessment is obvious – more narrative can mean more reading.
Funders already rely heavily on researchers giving their time to peer review. Simply replacing concise metrics with thousands of additional words would not necessarily improve assessment.
NHMRC’s evaluation of its Top 10 in 10 policy is useful here. Most surveyed reviewers believed the policy helped shift attention towards quality rather than quantity, and many reported reduced burden, but 23% did not consider burden to have decreased. Some reviewers noted that more detailed assessment of selected publications required careful consideration and time. But this does not mean the choice is between short applications and adequate evidence.
Possible approaches to providing reviewers with sufficient evidence while reducing unnecessary information could include:
- asking applicants to select fewer achievements rather than catalogue everything they have produced;
- using structured prompts so reviewers know where to find particular information;
- separating concepts that require different judgements;
- removing repeated information across application sections;
- asking only for evidence relevant to the assessment criteria;
- limiting supporting material that reviewers are not expected to use; and
- designing forms around the questions reviewers genuinely need to answer.
The SNSF model illustrates the option of fewer major achievements, described in more depth. UKRI illustrates another with structured modules covering different forms of contribution. Research Ireland separates a five-page track record narrative from the research programme and impact statement.
RoRI’s research on narrative CV implementation is also relevant here. It examines how narrative CVs change evaluative practice, including the ways reviewers interpret criteria, the role of panel chairs and the conditions that help or hinder the intended use of narrative information.
Reviewer burden should be reduced through better information architecture, not by removing information necessary for judgement. A short application that forces reviewers to guess is not necessarily efficient.
Responsible research assessment needs a common evidence base
When an application does not contain enough information to assess the criteria, reviewers may feel compelled to look elsewhere. They may turn to Google Scholar, Scopus, researcher profiles, journal websites, institutional biographies, or even the publications themselves.
While this may feel like a helpful way to fill gaps, it introduces inconsistency. One reviewer may conduct extensive external searches, another may only look things up when uncertain, and another may rely solely on the material provided in the application. As a result, applicants are no longer being assessed against a common evidence base.
External searching can also quietly reintroduce exactly the kinds of information that responsible research assessment is trying to move away from, such as h-indices, citation counts, and journal-level metrics.
At the heart of this issue is a crucial distinction between verification and evidence generation. It is entirely reasonable for a reviewer to verify claims that an applicant has already made. But it is a very different matter if the reviewer must find the evidence needed to construct the applicant’s case in the first place.
NHMRC makes this distinction in its research impact assessment. Reviewers may verify evidence supplied by applicants, but they are not expected to seek evidence to support claims that applicants have failed to evidence themselves.
UKRI’s Funding Service reviewer guidance goes further, saying that reviews must be based only on information contained in the application. It must also avoid using journal metrics, journal hierarchies, conference rankings, h-index, or i10-index as substitutes for assessing research quality. Reviews that fail these requirements can be treated as unusable.
Research Ireland’s Stage 2 Handbook takes a similar position. Hyperlinks cannot be used to provide information necessary for review or to circumvent page limits, and reviewers are not obliged to access linked material.
It is clear that responsible assessment requires a sufficiently complete, shared evidence base. Reviewers should assess the application, not reconstruct it.
Implementing responsible research assessment requires rules that work
There is also a problem with simply telling applicants that particular metrics are “discouraged”. If one applicant follows that instruction and another provides Journal Impact Factors, citation counts or other familiar indicators anyway, what happens? If nothing happens, the instruction risks becoming optional. And once a reviewer has seen the information, we cannot reasonably expect it to have no influence. You cannot unsee it!
This can create a collective-action problem for applicants. Researchers may worry that excluding a metric places them at a disadvantage if competitors include it and some reviewers continue to value it.
Research Ireland’s Investigators Programme provides one of the clearest current examples of a funder addressing this problem through application rules. Its prescribed narrative CV cannot include journal or publication metrics, total publication counts, or a long list of researcher-level indicators, including h-index-type measures. Hyperlinks and URLs are also excluded from the narrative CV. Failure to follow these requirements can make an application ineligible and that is a strong consequence.
Whether ineligibility is proportionate in every scheme is a separate question. Other systems might remove inadmissible material, require correction before review or prevent entry of certain information through application-system design. A central implementation challenge is designing a mechanism that keeps excluded information from influencing assessment.
UKRI addresses the reviewer side of this issue. Reviews must meet stated usability requirements, including avoiding journal and researcher metrics as substitutes for assessing quality. Reviews that fail those requirements can be removed from the assessment process.
This moves responsible assessment from aspiration towards process integrity, where rules must operate on both sides of the assessment system.
Training reviewers for responsible research assessment
Training also needs to go beyond telling reviewers what not to do. An instruction such as “Do not use the Journal Impact Factor” is not, by itself, training. If reviewers are asked to make nuanced qualitative judgements, a shared understanding of what those judgements involve needs to be developed. Reviewers may need to distinguish:
- methodological quality from journal reputation;
- scholarly attention from research quality;
- significance from citation accumulation;
- team achievement from applicant contribution;
- scholarly influence from influence beyond academia;
- evidence from promotional language; and
- contextual information from evidence of achievement.
Worked examples could help by illustrating what strong evidence of research quality looks like in practice. These type of resources help demonstrate what makes an explanation of applicant contribution convincing and how reviewers should respond when an applicant provides an inappropriate metric. They can also show how evidence should be interpreted differently across disciplines, how career context should shape judgement, and what a high-quality reviewer rationale looks like in a completed assessment.
CoARA’s qualitative assessor training resource suggests evaluating training through assessor confidence and consistency and notes the value funders place on reviewer exchange and peer-to-peer learning.
The ERC’s approach to researcher assessment directs reviewers away from publication counts and Journal Impact Factors towards qualitative assessment of scientific content. Applicants may select up to ten research outputs and explain their significance, their role in producing them and their relevance to the proposed project. Evaluation panels consider research field, career stage and personal circumstances.
The ERC has also described why implementation support matters. When it revised its assessment approach, the Scientific Council said it was preparing guidance for applicants and evaluators so they would understand the purpose of the changes, and that it would monitor the effects and refine the process based on experience and feedback.
But implementation ultimately depends on how reviewers apply these principles in practice.
That is why calibration matters as much as instruction.
Why reviewer discussion and panels can add value
Training can help reviewers before assessment begins. There is another potential opportunity for calibration while assessment is taking place, when reviewers are actively engaging with real applications and encountering the practical challenges of applying criteria in context.
This is one of the potential benefits of panel-based review, beyond aggregation of multiple independent scores. The value of a panel is not simply that more reviewers contribute to the decision.
Its value can lie in making individual judgements open to question, within a structured space where interpretations are compared and reasoning is discussed. A reviewer may interpret a criterion differently from another reviewer. They may give weight to a familiar indicator, interpret research quality through journal reputation or apply a scoring descriptor differently from colleagues.
When assessments are conducted independently and combined without discussion, those differences may never become visible. Deliberation can create an opportunity to ask why.
Why does this evidence demonstrate research quality?
Why has this application been assessed as outstanding rather than very good?
Are we assessing the applicant’s contribution or the achievement of the wider team?
Are we treating scholarly influence as evidence of research quality?
Have we interpreted the assessment criterion consistently?
DORA addresses this issue in its resource on Debiasing Committee Composition and Deliberative Processes. The guidance is relevant to funding panels as well as hiring and promotion committees. It focuses on bringing multiple perspectives into decision-making, fostering diversity of opinion, building transparency and reducing reliance on inherited decision-making norms.
CoARA’s commitments also place peer review at the centre of qualitative assessment and call for transparent processes, training and access, where possible, to reviews or deliberation outcomes.
Empirical research shows that discussion can influence grant assessment. In Pier and colleagues’ study of constructed NIH peer-review panels, reviewers within individual panels converged more after discussion, and the researchers identified “score calibration talk” as an important part of how scoring judgements were negotiated. However, agreement between panels did not improve, showing that deliberation should not be assumed to reduce variability.
An earlier analysis of NIH R01 reviews found that panel discussion had a practically important effect on more than 13% of applications. Other research has found that reviewers often value panel discussion, while also identifying limitations including uneven participation, time pressure and the role of panel chairs.
Evidence also cautions against assuming that panels automatically improve reliability. A study comparing two panels assessing the same 65 medical research proposals found that panel discussion did not improve inter-panel reliability compared with the mean of independent reviewer scores.
This is important because it argues against treating panels as inherently better. Panels can introduce their own risks: anchoring, dominant voices, group conformity, uneven participation and inconsistent chairing.
I do not recommend that all schemes include panels. Responsible assessment does, however, require a process that lets reviewers test, challenge and calibrate their judgements against both the criteria and others’ interpretations. For schemes without panels, that function must be designed elsewhere through strong pre-review calibration, worked examples, benchmark assessments, opportunities for reviewer exchange, systematic examination of reviewer rationales, and feedback after assessment.
The more a system relies on expert qualitative judgement, the more attention it may need to give to the processes that test that judgement.
Reviewer rationales are implementation data
Reviewer comments are an underused source of evidence about assessment reform. Reviewer rationales are usually treated primarily as feedback for the applicant. But they can also tell a funder whether the assessment system is functioning as intended.
Reviewer comments can show whether reviewers:
- understand what is being assessed;
- continue to rely on discouraged metrics;
- confuse scholarly influence with research quality;
- understand applicant contribution;
- apply career-context provisions consistently;
- use the score descriptors in the intended way;
- require information that the application does not provide; or
- interpret criteria very differently from one another.
NHMRC’s Responsible Research Assessment Action Plan provides a useful example of reviewer rationales as evidence about whether assessment reform is working. Its planned actions include reviewing peer-reviewer feedback for alignment with reform commitments, evaluating support and training for applicants and reviewers, examining how often applicants continue to provide Journal Impact Factors despite guidance against them, and evaluating the use of quantitative indicators in applications.
This closely aligns with CoARA’s commitment to evaluate assessment practices, criteria, and tools using evidence and research on research.
For funders using panels, another valuable source of learning may be recurring areas of disagreement, such as why reviewers are interpreting a criterion differently or why scores are moving after discussion.
These patterns can reveal where guidance, application design, reviewer training or scoring descriptors need improvement. This can be seen as both assessment quality assurance and organisational learning.
Implementing responsible research assessment means examining behaviour
A funder needs a way to know whether its reform has worked. Counting how many reviewers completed training is not enough. Introducing a narrative CV, revising guidance, or signing DORA are not enough either. These are reform activities, not evidence that assessment behaviour has changed. Funders could instead examine questions such as:
- Are reviewers still referring to journal reputation?
- Are citation indicators being used as evidence of research quality?
- Are reviewers searching externally for h-indices or other metrics?
- Do reviewers distinguish applicant contribution from team achievement?
- Are different career paths being recognised?
- Are applicants able to provide sufficient evidence against each criterion?
- Are reviewer rationales aligned with the published assessment criteria?
- Is scoring becoming more or less variable?
- Where do reviewers disagree most?
- Are reforms increasing or reducing reviewer workload?
- Do particular disciplines or career stages benefit or are disadvantaged by the new format?
- Are funding outcomes changing in ways the funder did not anticipate?
DORA’s 2026 Practical Guide for Research Funding Organizations and the accompanying Global Research Council self-assessment tool treat implementation as an ongoing organisational process. The self-assessment approach is intended to help funders identify gaps, assess their position across multiple dimensions of responsible assessment and plan further improvement.
This matters because responsible research assessment is, in large part, a behaviour-change intervention. NHMRC has itself mapped its Action Plan to a behaviour-change model, linking reform to changes in what is possible, easy, expected, rewarded and required within the assessment system.
Implementing responsible research assessment may require experimentation
We should not assume that the traditional architecture of peer review itself is fixed.
The CoARA framework for improving proposal assessment encourages funders to consider a range of established and emerging assessment practices and to adapt approaches to their own contexts rather than assuming one model will work everywhere.
The Research on Research Institute’s AFIRE programme is exploring alternative approaches to research funding assessment, including distributed peer review and experimentation with grant-allocation processes. RoRI’s wider work has also examined partial randomisation, where random selection is used at particular points among proposals already judged suitable for funding, rather than replacing assessment with a simple lottery.
RoRI has also developed guidance for distributed peer review, where applicants to a funding call participate in reviewing other applications. Funders have undertaken trials and implementations, providing another example of assessment architecture being treated as something that can be tested rather than inherited unchanged.
I do not think this means that lotteries or alternative review models should become requirements of responsible research assessment.However, it does suggest a broader principle, one which many researchers I have worked with would like to see implemented.
Funders should be willing to ask whether the assessment process they have inherited remains the best mechanism for making the decision they need to make. For example, if a group of applications all meet a high funding threshold, can peer review reliably distinguish the 23rd-best proposal from the 27th-best proposal? Or are we asking reviewers to manufacture a level of precision that the assessment process cannot support?
Research on grant review supports this question. An Australian analysis of NHMRC panel scores found substantial uncertainty around funding outcomes when variation between reviewers was considered, illustrating how fine distinctions between applications can be sensitive to the reviewers assigned to them. (Graves, Barnett and Clarke, 2011).
Questions such as these deserve experimentation rather than assumption.
To implement responsible research assessment, the key point is that the assessment method itself should remain open to evaluation and improvement.
Seven questions for funders implementing responsible research assessment
Across these examples, I think there are seven practical questions that funding organisations can ask of any assessment reform (See these outlined in the image below).
From declaration to decision
There is considerable momentum behind responsible research assessment. But the next phase of reform is harder than signing declarations, revising guidance, or redesigning CV templates. It requires attention to the architecture of assessment itself and to the conditions under which reviewers are expected to make judgements.
If funders want reviewers to rely less on traditional metrics, they need to provide credible alternatives and enough information to use them. If they want more qualitative judgement, applicants need sufficient space to provide meaningful evidence of quality, significance and contribution. At the same time, reviewer burden needs to remain manageable, which means removing unnecessary information rather than restricting information essential to the judgement being asked for.
The same principle applies to the evidence base. If funders do not want reviewers to rely on external metrics or information gathered outside the application, then applications need to contain enough relevant evidence to support the assessment. If particular indicators should not influence funding decisions, their exclusion cannot depend entirely on applicants voluntarily omitting them or reviewers attempting to disregard them after they have already been seen. The rules, application design and reviewer guidance all need to reinforce the same assessment intent.
As qualitative judgement becomes more central, reviewer capability also becomes more important. Training needs to help reviewers understand the criteria, the kinds of evidence that support them and the distinctions between research quality, scholarly influence, applicant contribution and wider impact. Opportunities for calibration and deliberation can add another layer of quality assurance by allowing individual judgements to be tested, questioned and contextualised.
Funders also need to know whether these changes are influencing assessment practice. That means looking beyond whether reviewers completed training or introduced a new narrative CV. Reviewer rationale, patterns of disagreement, use of inappropriate indicators, external searching, workload, and funding outcomes can all provide information about whether the assessment system is functioning as intended.
The principles behind responsible research assessment are becoming increasingly well established. The implementation challenge is now much more practical. It comes down to whether the assessment system has been designed so reviewers can make the judgement it asks of them. That may be the more meaningful test of responsible research assessment.
