
Signing DORA is not enough: responsible research assessment must reach the point of judgement
On 24 June 2026, the Australian Research Council (ARC) published its Commitment to Responsible Research Assessment and signed the San Francisco Declaration on Research Assessment, commonly known as DORA. The ARC statement recognises that traditional metrics alone cannot capture the full value of research and commits the agency to fair, inclusive, transparent and contextually sensitive assessment. Its principles apply to National Competitive Grants Program assessments and evaluations, and to the Research Insights Capability [1].
This is an important and timely development for Australian research.
Other organisations have recognised the need for their leadership in this area. The National Health and Medical Research Council (NHMRC) endorsed DORA in 2013, introduced its Research Impact Track Record Assessment Framework in 2018, implemented the Top 10 in 10 publications approach in 2022, joined the Coalition for Advancing Research Assessment, or CoARA, in 2024, and published a 2025-2030 Action Plan for Responsible Research Assessment [2].
I welcome these developments. Both funders have taken meaningful steps to broaden what is recognised and change the information used in assessment. My concern is about the implementation of these reforms. Changing an application form is not the same as changing assessment practice and signing a declaration is not the same as implementing it. Even revised criteria may have limited effect if the people applying them continue to rely on the indicators and assumptions that have shaped academic assessment for decades.
Responsible research assessment must ultimately reach the point at which an individual reviewer makes a judgement.
What happens when the criteria change but the judgement does not?
In my work supporting researchers to identify and communicate their research contributions and impact, I regularly see how formal criteria are interpreted through reviewer feedback.
I have heard researchers involved in peer review say that they continue to look primarily at publication records when deciding whether an applicant has a strong track record. I have read feedback that places considerable weight on citation counts, journal standing or the perceived prestige of selected publications. I have also encountered questions about why a paper has relatively few citations, even when it is recent, highly specialised, focused on a rare disease or small population, or has generated influence through policy, guidelines or professional practice rather than extensive academic citation.
These observations come from direct conversations with researchers, written reviewer feedback and informal workshop discussions. They are not a systematic study of peer-review behaviour, and I do not suggest they represent all reviewers. Many reviewers make a serious effort to understand the changes and apply revised criteria fairly. However, the persistence of these examples raises an important question: are we changing the assessment system, or simply placing new forms around old assessment habits?
When research assessment criteria change but reviewer judgement does not
Researchers whose contributions are least visible through traditional publication and citation measures are most likely to be disadvantaged. This may include researchers working in small or less citation-dense fields, those addressing rare diseases or locally important problems, early-career researchers, people with interrupted or non-linear careers, and researchers whose strongest contributions have been made through teams, infrastructure, data, methods, policy, professional practice or engagement. These are among the diverse contributions and career pathways that CoARA seeks to recognise.
Researchers and institutions that adopt responsible assessment guidance may also fear that they are placing themselves at a disadvantage. If some applicants continue to include h-indices, journal rankings, citation counts or field-weighted indicators, and some reviewers continue to value them, applicants may think that leaving it out is a risk. The concern becomes: if a reviewer prefers those measures and other applicants provide them, will I disadvantage myself by strictly following the guidance?
This means the people who follow the reform most closely may end up feeling the least protected. It can also push institutions to keep using discouraged indicators, since they can’t be sure every other institution will interpret the guidance the same way.
Funders also lose when the judgements made through peer review do not align with the qualities their assessment criteria claim to value. Relying on prestige proxies risks rewarding familiarity over researchers who genuinely meet the criteria for quality, contribution, leadership and impact.
Reviewers also lose as they are asked to make more nuanced and contextual judgements without always being given a shared understanding of what should replace the traditional measures. When criteria remain open to different interpretations, reviewers may return to familiar numbers because they appear to offer a more defensible basis for differentiation.
Ultimately, the research system and the public both lose. Research assessment has the power to shape what researchers pursue, document and prioritise. The immediate loss falls on the individual applicant. The deeper loss falls on the funder, whose assessment process no longer identifies what it claims to value. The wider loss falls on the research system itself when reform exists in policy but not in judgement.
The people implementing the reform may not know it exists
There is another aspect of the implementation gap that I encounter regularly. When I mention DORA to researchers, including those who undertake peer review, many tell me they have never heard of it. They may not know that the NHMRC endorsed DORA, joined CoARA, or that changes to track record assessment sit within a wider international movement to reform how research and researchers are assessed.
These are practice-based observations, not formal findings. I cannot know what information every reviewer has received. Nor does unfamiliarity with DORA necessarily mean that a reviewer is assessing irresponsibly. Someone may apply its principles without knowing the declaration, while knowing the declaration does not guarantee that it will shape their judgement.
Nevertheless, the lack of awareness points to a deeper issue. When criteria change without a clear explanation of the problem they aim to address, reviewers may interpret them as administrative redesign rather than cultural reform. Selecting ten publications may be understood simply as providing a shorter publication list or reducing burden. A narrative track record may be treated as extra writing rather than a shift away from publication volume, journal prestige and citation proxies. An impact claim may be seen as supplementary to the “real” track record rather than evidence of a different form of contribution.
Without understanding why the system has changed, reviewers may reconstruct the old system inside the new one. They may search for an applicant’s h-index, total publications, journal rankings or Field-Weighted Citation Impact because those indicators still feel like the most credible basis for comparison.
Implementation therefore requires more than instructions for completing a review. Funders need to communicate the rationale for reform across the sector. Reviewers need to understand the commitments behind revised criteria and the limitations of familiar indicators. They also must learn about the types of evidence that can support contextual judgements of quality, contribution and impact. Researchers need this understanding too, because they are applicants, mentors, institutional leaders and often peer reviewers themselves.
Why traditional research metrics remain attractive
Publication counts, citation numbers, Journal Impact Factors, h-indices and field-normalised citation indicators offer very attractive assessment systems: a quick way to reduce a complex research career to a small number of apparently comparable measures.
Reviewers working under pressure can use those numbers to sort, rank and differentiate applicants. Numbers may also appear more objective and defensible than qualitative judgements.
The problem is not that all quantitative indicators are useless. It arises when an indicator is treated as evidence of something it does not measure. DORA’s guidance emphasises that no indicator can capture research quality in one number and recommends that organisations be clear, transparent, specific, contextual and fair in their use [3].
The relevant question is not simply, “Is this a good metric?” It is: what are we assessing, and does this indicator provide relevant evidence for that judgement?
Journal prestige is not evidence of research quality
The Journal Impact Factor (JIF) is a clear example of an indicator escaping its appropriate unit of analysis. It describes average citation performance at journal level. It does not assess the rigour, originality, significance or usefulness of an individual article. Citation distributions within journals are also highly skewed, so a journal average tells us little about a particular paper.
DORA recommends that journal-based indicators should not be used as surrogate measures of the quality of individual articles or researchers’ contributions. It asks funding agencies to make criteria clear and emphasise that scientific content matters more than publication metrics or the identity of the journal [4].
Yet journal names continue to be powerful signals of quality. A reviewer may never write down the Impact Factor but may still assume that publication in a prestigious journal indicates excellent research. Conversely, rigorous work in a specialist, regional or less prestigious journal may be undervalued before its content is considered.
Removing a metric from an application does not necessarily remove the status hierarchy it helped create.
Citations show scholarly attention, not research quality
Article-level citations are more closely connected to a publication, but they also have limitations. Citations accumulate slowly, differ among fields and publication types, and are affected by the size and pace of the research community. A paper addressing a common condition or rapidly expanding field has a larger potential citing audience than work on a rare disease, local policy issue or emerging area.
Citations may reflect scholarly influence, but influence does not necessarily indicate quality. A paper may be cited because it produced an important finding, introduced a method, attracted controversy, contained an error or became something others wished to challenge.
Citations can provide evidence of scholarly attention. They do not, on their own, establish research quality, policy influence, clinical uptake, technological deployment or societal benefit.
The h-index in responsible research assessment
The h-index combines publication volume and citation accumulation. A researcher has an h-index of 20 when at least 20 papers have received at least 20 citations. It is easy to understand and compare. Those qualities have contributed to both its influence and its misuse.
Jorge Hirsch’s original 2005 paper proposed the h-index as a single indicator for characterising scientific output and supporting comparisons in recruitment, promotion and grant allocation [5]. However, Hirsch acknowledged that one number could offer only a rough approximation of a multifaceted profile. Values differ between fields; non-mainstream areas may not accumulate the same citations; the measure can undervalue researchers with a few seminal papers; and large collaborations can obscure individual contribution [5].
DORA identifies further concerns. The h-index depends on the database used, rises with career length, can disadvantage people with career interruptions, varies by discipline and says little about an author’s contribution [3].
The h-index does not directly measure research quality, nor technological, policy, clinical or societal impact. That is not a failure of the indicator, since it was not designed to measure those things. The failure occurs when an assessment system uses it as though it does.
Does FWCI support responsible research assessment?
Field-Weighted Citation Impact, or FWCI, is more sophisticated than a raw citation count. It adjusts for differences among fields, publication types and years. In simplified terms, an FWCI of 1 indicates approximately the expected citation performance for comparable work, while 2 indicates roughly twice the expected performance.
This makes FWCI more informative than comparing raw citations across unrelated fields. It still does not transform citation performance into research quality. Results depend on field definitions and publication classification, highly cited outliers can distort small publication sets, and values change as citations accumulate.
DORA recommends applying field-normalised indicators to large datasets rather than individual researchers, where small samples, outliers and fluctuations undermine reliability [3]. FWCI is not inherently incompatible with responsible assessment, but using an applicant’s FWCI as shorthand for research quality would be difficult to reconcile with that guidance
We need to be clearer about what we are assessing
One of the problems in research assessment is that research quality, contribution, scholarly influence, influence beyond academia and applicant contribution are often treated as though they are the same thing. They are not. Each involves different questions and requires different kinds of evidence.
This distinction matters because the same piece of evidence cannot do every evaluative job. A policy citation may show that research entered a policy process, but it does not automatically establish methodological quality or policy change. A patent may indicate a connection to invention, but not adoption or societal benefit. A high citation count may suggest scholarly attention but not impact beyond academia. Responsible assessment should therefore not replace one universal measure with another. The task is to identify what is being assessed, understand its context, and examine evidence that is appropriate to that judgement.
Research quality, contribution, scholarly influence and research impact are often used interchangeably. They shouldn’t be. Research quality may include methodological rigour, validity, originality, reproducibility and ethical conduct. Contribution to knowledge may include resolving a question, changing understanding, introducing a method or enabling further investigation. Scholarly influence may include citations, replication and reuse. Influence beyond academia may include policy, clinical guidelines, professional standards, industry practice, technology, environmental management or community decision-making. Applicant contribution concerns what the individual did, especially within collaborative research.
Each requires different questions and evidence. A policy citation may show that research entered a policy process, but it does not automatically establish quality or policy change. A patent may demonstrate a connection to invention but not adoption. A high citation count may indicate scholarly attention but not societal benefit.
Responsible assessment should not replace one universal measure with another. CoARA calls for recognition of diverse outputs, practices and activities, with assessment based primarily on qualitative judgement and supported by responsible use of indicators [6]. The task is to identify the contribution being assessed, understand its context and examine appropriate evidence.
What DORA and CoARA require from responsible research assessment
DORA emerged from concern about the misuse of journal-based indicators, particularly the Journal Impact Factor, in funding, appointment and promotion. Its recommendations extend beyond removing journal metrics. DORA asks funders and institutions to assess research on its merits, recognise outputs beyond articles, make criteria explicit and consider qualitative evidence of influence on policy and practice [4].
CoARA takes the reform further. Its Agreement calls for diverse contributions and careers to be recognised, qualitative evaluation with peer review at its centre, responsible use of indicators, and removal of inappropriate uses of the Journal Impact Factor and h-index [6].
Notably, CoARA does not treat signing as implementation. Its commitments include allocating resources, revising criteria and tools, raising awareness, providing guidance and training, communicating progress and evaluating whether reforms work [2,6]. The implementation concern I am raising is therefore embedded within the principles themselves.
Responsible research assessment in Australia: the NHMRC experience
The NHMRC provides a useful Australian example because its reforms have generated evaluation evidence. Its Top 10 in 10 policy was intended to shift attention from publication quantity towards the quality of selected research and the applicant’s contribution [7].
The evaluation reported encouraging results. Most surveyed reviewers supported the policy and agreed that it helped emphasise quality rather than quantity. It reduced emphasis on total publication numbers, Journal Impact Factors and similar indicators, and reduced burden for many reviewers [7].
But older habits persisted, as some reviewers sought full publication lists or independently searched for FWCI, publication totals and h-indices. Others requested first- or last-author counts, Q1 publications and other metrics. The report concluded that more detailed guidance was needed for applicants and reviewers [7].
This does not mean the policy failed, rather it shows that changing the information presented does not automatically change what reviewers believe they need for a credible judgement. When familiar indicators are removed, some may search for them elsewhere.
The NHMRC has recognised this challenge. Its Action Plan includes reviewing reviewer feedback for alignment with CoARA, evaluating training and support, reinforcing guidance on publication-based metrics and communicating reform across the sector. It maps actions to a model of behaviour and culture change, acknowledging that behaviour is shaped by skills, norms, incentives and policy [2].
Why this is not a criticism of individual reviewers
Reviewers have been trained, rewarded and promoted within systems that privileged publication volume, journal prestige and citations. They are now asked to make more contextual and qualitative judgements, often while reviewing complex applications under considerable time pressure and, by all accounts, without sufficient communication of changes and training in new assessment criteria.
It is far easier to recognise a journal’s name, check an author’s listed position, or tally citations than it is to genuinely assess a paper’s rigour and significance, an individual’s contribution, or the influence of research on policy and practice.
Reviewers cannot be expected to infer the philosophy of responsible assessment from a revised form. Without an explanation of DORA, CoARA and the rationale behind new criteria, they may interpret the task using the norms that have guided assessment throughout their careers.
Responsible assessment must therefore be treated as a systems implementation challenge, not a compliance exercise. Funders and institutions need to know whether reviewers understand why criteria changed, whether rationales still rely on inappropriate proxies, and whether specialised, emerging or locally important work is being disadvantaged. Most importantly, they need to know whether applicants are being assessed differently.
The ARC’s opportunity to implement responsible research assessment
The ARC’s new commitment recognises diverse outputs and academic activities, varied approaches to research translation, disciplinary and cultural context, and both quantitative and qualitative information. It also commits to guidance on the use and limitations of information generated through evaluation [1].
It is too early to judge implementation, which creates an opportunity to build it into the reform from the beginning. Success should not be measured only by whether principles appear in policies or assessors receive another guidance document. It should include whether those principles are evident in assessment reasoning, panel discussions, scoring patterns and feedback.
From commitment to practice
DORA and CoARA are not asking us to abandon judgement or eliminate all quantitative information. They ask us to use evidence for purposes it can reasonably support, make qualitative expert assessment central, and recognise diverse contributions, careers and pathways to influence. I acknowledge that this is a more demanding process.
It requires reviewers to look beyond convenient proxies and funders to support them. It requires clear criteria, appropriate examples, reviewer preparation, shared interpretation, calibration and evaluation. It also requires institutions and researchers to stop reinforcing the prestige markers and assessment habits that these reforms are intended to move beyond.
The central question is no longer whether an organisation has signed DORA or CoARA. It is whether the principles have reached the point at which a reviewer decides what counts as quality, contribution and impact. New forms and guidance are both steps forward however responsible research assessment will only be implemented when the judgements change too.
The real test is whether a researcher can be confident that their work will be assessed for what it has contributed, in context, rather than through the prestige metrics surrounding it.
I’ll be writing more soon about what that will actually take.
Looking for support with your next grant application? Explore our grant workshops and consulting services.
