Research evaluation feels precise.
Dashboards display citation counts to the nearest unit. Rankings separate institutions by marginal numerical differences. Bibliometric indicators are reported with decimal places. Comparative charts imply objectivity.
The visual language of evaluation conveys certainty.
However, research evaluation data are often more fragile than they seem.
This fragility arises from complexity, commercial opacity and the historical evolution of indexing systems that were never designed for how they are now used.
Universities are complex organisations. Scholars collaborate across borders. Affiliations change. Names are ambiguous. Databases are compiled at scale. Algorithms attempt to resolve identities and categorise outputs. Small inconsistencies compound over time.
When evaluation systems rely heavily on such data, fragility matters.
Let’s look at this issue through two lenses.
The Governance Lens: Precision and Illusion
From a governance perspective, research evaluation data serve essential functions. They inform strategy. They support benchmarking. They underpin performance frameworks. They provide evidence to regulators and funders. Perhaps, most importantly, the give a feeling of confidence. They make the board feel comfortable, giving them the a feeling that they are in control and they know what is happening.
Boards require clarity. Executives require comparability. Decisions cannot rely solely on anecdote.
But the path from scholarly activity to institutional dashboard is not straightforward.
Author disambiguation remains imperfect. Common names must be matched accurately across thousands of publications. Minor inconsistencies in spelling or affiliation can split a record or merge it incorrectly. Even leading bibliometric platforms continue to struggle.
Affiliation data are equally complex. Scholars hold multiple appointments. They move institutions. They publish during transitions. Attribution rules differ across ranking systems. What benefits one framework may disadvantage another.
Citation data vary by database. Coverage differs across disciplines. Update cycles are inconsistent. A citation count appears definitive, but it reflects indexing choices rather than an absolute reality.
When indicators are reported to two decimal places, they acquire authority. When institutions are ranked sequentially, marginal differences appear meaningful. When dashboards show incremental improvement, it is tempting to assume structural strength.
The distortions are rarely dramatic. They are incremental.
Yet when metrics are aggregated, benchmarked and linked to resource allocation, small distortions can influence large decisions.
Through a governance lens, the central question is this:
How robust are our decisions when they are based on the imperfections in data that inevitably exist?
The Scholar’s Lens: Evaluation and Identity
Promotion panels review citation counts. Funding bodies assess track record. External reviewers consult database profiles.
When data misrepresent a scholar’s record, consequences follow.
An author whose outputs are split across multiple profiles may appear less productive. An academic incorrectly merged with another may appear more productive, until the error is discovered. Inaccurate affiliations may obscure institutional contribution.
Even when broadly accurate, interpretation can be superficial. Citation patterns differ markedly between disciplines. Journal quartiles are sometimes treated as shorthand for quality without context. In a recent paper, I examined transparency concerns surrounding the reporting of h-indexes, Journal Impact Factors and CiteScores (Kendall, 2024).
Researchers often assume evaluation systems see their work clearly.
In reality, what is seen is a structured representation filtered through database design, indexing rules and algorithmic matching.
Recognising this is not an invitation to manipulate profiles. It is a reminder that structured data require attention.
Maintaining accurate publication records and consolidated scholarly identities is not administrative housekeeping. It is strategic maintenance.
Aggregation and Amplification
Fragility increases with aggregation.
Individual outputs become departmental summaries. Departments become institutional dashboards. Institutions feed national assessments and international rankings.
At each stage, simplification occurs and errors can be introduced.
A department’s intellectual character cannot be reduced to average citation impact. An institution’s research culture cannot be captured fully in a composite index. Yet aggregated figures become proxies for quality. Institutions report those figures, the general public accept them as hard evidence and students and scholars based decisions on them.
Over time, those proxies influence hiring, student recruitment, resource allocation and strategic direction.
None of this is irrational.
But it rests on data that are constructed, filtered and, at times, incomplete.
Fragility does not invalidate evaluation. It complicates it.
The Cultural Consequence
When evaluation data are treated as definitive rather than indicative, cultural effects follow.
Leaders may chase marginal ranking improvements. Scholars may optimise visible indicators rather than invest in slower, riskier or interdisciplinary work.
At its core, this is a question of trust.
- Do leaders trust the data sufficiently to base major decisions upon them?
- Do scholars trust that evaluation systems represent their work fairly?
- If trust is partial, how should systems be designed to mitigate that uncertainty?
Toward More Robust Evaluation
The solution is neither abandonment of metrics nor blind faith in them.
It is maturity.
For governing bodies, this may mean triangulation: combining quantitative indicators with informed peer review and contextual judgement. It may involve auditing data quality and resisting the overinterpretation of marginal differences.
For scholars, it may mean maintaining accurate profiles, understanding how databases index work, and recognising how evaluation committees interpret bibliometric signals.
Questions Through Two Lenses
Through the Governance Lens
- How confident are we in the data underpinning our key research indicators?
- Do our dashboards encourage overinterpretation of marginal differences?
- Where might aggregation be masking variation?
- How do we combine metrics with informed judgement?
Through the Scholar’s Lens
- Is my publication record accurately represented across major databases?
- Are my affiliations consistently recorded?
- How might committees interpret the pattern of my work?
- Do I understand how the metrics I cite are constructed?
Continuing the Conversation
If your institution relies heavily on research evaluation indicators, this is not a technical issue. It is a governance issue.
Understanding how bibliometric systems are constructed, where they are robust, and where they are vulnerable, strengthens strategic decision-making. It also protects institutions from overinterpreting marginal differences that may rest on imperfect data.
I work with universities and academic leaders on research evaluation design, metric interpretation and governance alignment. If these questions resonate with your current discussions, I welcome a thoughtful conversation.
Reference
Kendall, G. (2024) More Transparency is Needed When Citing h-Indexes, Journal Impact Factors and CiteScores. In Publishing Research Quarterly, 40 (1): 80-99. http://dx.doi.org/10.1007/s12109-024-09983-3