RSS Amplifier

The Workforce Lens’s Substack · Jun 21, 2026

The Visibility Tax: Why Hybrid Performance Reviews Fail and How to Fix Them

0
Sign in to vote or save

Dominika Borna · The Workforce Lens’s Substack

📌In a Nutshell:

Hybrid work did not break performance evaluation; it exposed that performance evaluation was already broken

  1. The Career Stagnation Paradox: Why professional advancement often stalls even when individual productivity remains at an all-time high

  2. The Cost of Invisibility: How physical absence from the office translates into a tangible financial and professional tax on the future of an employee

  3. The Presence Delusion: Why organisations have spent a century maintaining the fiction that physical attendance is a reliable proxy for actual work

  4. The Developmental Deficit: How the loss of casual office interaction is not merely a social inconvenience but a structural failure that alters career trajectories

  5. The Responsiveness Trap: Why the pressure to answer messages immediately creates a facade of productivity that actively sabotages the quality of the work

  6. The Equity Failure: How the flexibility intended to support diverse talent is being weaponised by outdated systems to restrict their progress

  7. The Systemic Reality: Why the primary issue is not the individual manager but a century-old evaluation framework that is fundamentally unfit for the modern era

  1. The Hidden Conflict Between Hybrid Productivity and Career Growth

  2. Why Physical Presence Remains the Default Metric for Performance

  3. The Invisible Feedback Gap and Its Impact on Professional Development

  4. The Responsiveness Trap: How Performative Availability Sabotages Deep Work

  5. How Proximity Bias Reinforces Inequality and the Gender Promotion Gap

  6. Why Modern Organisations Struggle to Transition to Output-Based Evaluation

  7. Structural Solutions to Eliminate Proximity Bias and Build Evaluation Equity

  8. Final thoughts: Redefining Performance for a Post-Presence Era

  9. Key Takeaways

  10. Next on The Workforce Lens

  11. Further Reading

For most of the twentieth century, organisations solved a genuinely difficult problem by ignoring it. Measuring individual performance — what someone actually contributes, at what quality, with what effect on the people and work around them — is hard. It requires agreement on what matters, investment in observation, and tolerance for the ambiguity that comes with assessing human effort in complex environments. Presence was easier. If you could see someone working, you could tell yourself you knew how well they were working. The substitution was never announced. It did not need to be. It was simply the water organisations swam in, invisible precisely because it was everywhere.

Hybrid work drained the pool.

What remote and hybrid arrangements exposed is not a new problem but an old one, newly legible. The infrastructure that organisations use to evaluate people was never measuring what it claimed to measure. It was measuring proximity. The claim was performance. That distinction, suppressed for decades by the assumption that everyone would always be in the same building, is now impossible to avoid — and the evidence for it has accumulated to the point where dismissal requires more effort than acknowledgement.

At Trip.com, a randomised trial tracked 1,612 employees across two years — half shifted to hybrid schedules, half remaining fully in-office — in a study published in Nature in 2024 by Bloom, Han and Liang. The results were unambiguous on the productivity question: hybrid workers showed no meaningful difference in performance ratings or promotion outcomes over the study period. Attrition told the more consequential story. Quit rates fell by roughly a third overall; among non-managers specifically, from 7.2% down to 2.4%. Location, the study established, does not damage output.

If the productivity numbers are stable, the developmental reality is anything but.

A separate and growing body of research finds that fully remote workers face promotion penalties their actual output does not explain. The work is equivalent. The advancement is not. The gap does not close when performance is controlled for. It persists — not at the margins, but consistently, across industries and organisational types — in a pattern that points directly at the set of cognitive and organisational mechanisms that physical proximity quietly powered, and that distance just as quietly dismantles.

💡 Related read:
Remote work vs. productivity myths: what the data really says about performance, accountability and engagement

Discover what recent data reveals about remote and hybrid work, and explore how organisations can strengthen productivity, accountability, and employee engagement in flexible work models

Most organisations have never had a rigorous method for measuring individual performance. That is not a provocative claim — it is a structural reality that anyone who has sat through a performance calibration meeting will recognise immediately. What they have had, for most of the twentieth century and well into this one, is a workable substitute: physical presence. If a manager could see someone at their desk, in the meeting, in the corridor after the meeting, contributing visibly to the rhythm of the working day, that visibility quietly stood in for evidence. Nobody called it a proxy. It did not feel like one. It felt like knowing how someone was doing.

Hybrid work did not introduce a new distortion. It removed the conditions that made the old one invisible.

The availability heuristic — one of the most durable findings in behavioural psychology — describes a consistent pattern in human judgement: when people assess likelihood, quality, or performance, they draw disproportionately on whatever information is most mentally accessible, not whatever information is most accurate. For a manager with in-person reports, accessible information is dense, continuous, and largely ambient. Hallway exchanges. Body language across a conference table. The informal moment after a meeting where someone demonstrates exactly how they think under pressure. None of this is formally recorded. All of it accumulates into a felt sense of how a person operates — and that felt sense, when performance review season arrives, is extraordinarily difficult to separate from formal evaluation.

Remote workers generate a categorically different kind of signal. Scheduled calls. Submitted deliverables. Emails read asynchronously and out of context. The record is episodic rather than continuous, curated rather than ambient. An in-person employee leaves hundreds of small impressions across a year. A remote employee leaves a folder of outputs and a series of calendar invitations. Both may have done work of identical quality. One of them is substantially easier to remember doing it.

UCLA’s Breaking Bias initiative captures part of this through its SEEDS framework, which identifies distance bias as one of the brain’s habitual shortcuts — the tendency to assign greater weight and greater value to whatever feels immediately proximate. A 2021 study in Frontiers in Psychology extended this into the specific context of management, documenting how managers experience genuine uncertainty about what constitutes evidence of performance when direct observation disappears. Into that uncertainty, the research found, managers import substitutes: response times, calendar availability, the speed at which a message is answered. These are presence proxies. They measure accessibility. They are, in most professional roles, almost entirely unrelated to the quality of the work being done — and they migrate into performance assessments with a reliability that suggests the substitution is not occasional but structural.

Proximity does not only shape how performance is perceived. It shapes how performance develops. That distinction is where the argument becomes concrete — and considerably more uncomfortable for organisations that have treated the visibility question as a fairness issue rather than a measurement one.

Emanuel, Harrington and Pallais, in NBER Working Paper 31880, tracked code review comments among software engineers as a measure of developmental feedback. When offices were open, engineers co-located with their colleagues received 23.9% more comments on their code than engineers in separate buildings. The gap narrowed after offices closed and physical proximity was equalised by circumstance. Code review comments are not social pleasantries. They are the specific, iterative inputs through which technical skill develops — corrections, alternative approaches, explanations of why one solution holds up better than another under production conditions. A 23.9% difference in their frequency, compounded across months and years, does not produce a marginal difference in professional development. It produces a material one. And that material difference in development produces, in turn, a material difference in the performance record that managers eventually evaluate.

What organisations measure as a performance gap may, in a significant share of cases, actually be a feedback gap — a deficit in developmental investment that accumulated upstream, invisible in the moment, only legible in its effects at the point of promotion. Managers then assess that deficit as evidence of lesser capability. They are not being dishonest. They are reading accurately from a record that was shaped, long before the review conversation, by who sat near whom.

Cullen and Perez-Truglia, writing in the American Economic Review, close the loop. Using quasi-random variation in manager rotation at a large bank, they isolated the effect of face time on promotion outcomes independent of measured output. More face time with a manager produced higher promotion rates. At the firm they studied, this mechanism alone accounted for roughly a third of the gender gap in promotions. Not performance differences. Not qualification gaps. The simple fact of physical proximity to the person responsible for evaluating you.

The chain is exact. Visibility produces feedback. Feedback produces development. Development produces the performance record. The performance record produces the promotion. Remove visibility from the first link and the entire sequence degrades — not dramatically, not immediately, but persistently, in ways that only become fully visible when the career data is examined in aggregate.

Sarah is a Senior UX Designer. On a Tuesday in October, she finishes a redesign of a checkout flow that has been causing a 4% drop-off rate — fiddly, detail-intensive work that required three days of user research, two rounds of iteration, and a protracted negotiation with the engineering team about what was actually buildable in the available sprint. She submits it. She gets back a two-word acknowledgment. She waits.

She follows up two days later, framing the message carefully so it does not read as chasing. She updates her Slack status to “available” even though she is in the middle of something that requires concentration, because the status feels like a signal she cannot afford not to send. Her laptop produces the specific double-ping of a new Slack notification. She answers it in four minutes — not because four minutes is the right amount of time to give that message, but because she has absorbed, without anyone stating it explicitly, that response speed is being tracked somewhere in the background of how she is perceived. She attends a Thursday meeting with no agenda item requiring her presence, because not attending feels professionally riskier than an hour of her time.

None of this is productive. Every bit of it is rational.

The responsiveness trap operates below the level of policy. No manager has told Sarah that Slack response times factor into her performance rating. No document specifies that meeting attendance functions as a visibility signal. The behaviour emerges from an accurate reading of an environment where presence has become load-bearing — where being seen to be available substitutes for evidence that the work is good. Sarah has correctly identified that the evaluation system is measuring something other than her output and is optimising accordingly. The cost of that optimisation is paid in exactly the kind of deep, focused work that produced the checkout flow fix in the first place.

The performance review, when it arrives, is nominally about the quality of what she produced. In practice it is shaped, in ways her manager could not easily articulate, by the texture of her visibility across the year. The checkout flow redesign may surface in that conversation. The Slack status and the four-minute response time already have.

Proximity bias does not distribute its costs evenly. The Emanuel et al. research found that the proximity-driven feedback gap was larger for women than for men — an asymmetry that does not resolve under remote conditions but persists, leaving existing disadvantages structurally intact. Whether remote work actively amplifies gender gaps or simply fails to close them is a question the data does not fully resolve. The direction of the effect is consistent either way.

The Cullen and Perez-Truglia paper specifies the mechanism. The face-time advantage that drives promotion rates accrued to men and did not accrue equivalently to women. Same proximity. Different return on it. In hybrid settings where face time is unevenly distributed by default — determined partly by caregiving responsibilities, commute distance, and the social dynamics of who feels entitled to claim physical presence as a professional tool — that asymmetry does not disappear. It compounds.

The structural trap is at its most perverse in the design of flexible work itself. A Wharton faculty working paper from 2023 found that remote and flexible roles attract more diverse applicant pools. Flexibility opens doors that rigid in-person structures kept closed. The people who walk through those doors then encounter a promotion infrastructure built around the assumption of physical presence. Flexible work widens entry and narrows advancement simultaneously — and the populations most drawn to it, for entirely rational reasons rooted in how workplaces have historically been structured, are the most exposed to that narrowing. This is an argument from synthesis across several bodies of work rather than a single clean finding, but the data points pull consistently in the same direction.

In 2024, Dell formalised what most organisations practice informally. Remote employees were classified as ineligible for promotion by stated policy. Close to half of Dell’s workforce — figures drawn from contemporaneous reporting rather than primary disclosure, worth verifying against primary sources before publication — chose to remain remote regardless. That is not confusion about the terms. It is a revealed preference: workers had priced the penalty, weighed it against the value of flexibility, and decided to absorb the career cost. The calculation is rational. The system that makes it necessary is not.

Some employer surveys from 2024 and 2025 have flagged office attendance as an emerging explicit factor in performance evaluation, though the evidence base is inconsistent and the surveys vary considerably in methodology. The directional shift is clear enough: proximity bias, which spent decades operating informally and deniably, is in some organisations becoming codified. Codification has a certain honesty to it. It also creates legal and ethical exposure that informal practice did not, because once a criterion is written into policy, it can be challenged, audited, and litigated.

The reason organisations have not fixed this is less dramatic than most accounts suggest. Building evaluation infrastructure around outputs requires reaching explicit, documented agreement on what good work looks like in each role, what outputs constitute evidence of it, and how those outputs should be assessed at what frequency. Most organisations have never done this. What they have instead is a set of inherited performance management rituals — annual reviews, competency frameworks, calibration meetings — designed around physical presence and not meaningfully updated since that assumption stopped being universal. Presence is a poor measure of performance. It is, however, a legible one. Output-based measurement forces questions into the open that presence-based evaluation quietly suppresses. Most organisations prefer the quiet.

Deloitte’s 2024 Global Human Capital Trends report notes that existing metrics struggle to capture the complexity of modern work — a carefully hedged observation from a survey instrument, but one that reflects something real: organisations are broadly aware their measurement systems are inadequate, and most are not doing much about it.

Buried in the Bloom, Han and Liang study is a finding that receives far less attention than the attrition numbers. At the outset of the Trip.com trial, managers predicted hybrid work would reduce productivity by 2.6%. By the end, the same managers believed it had improved by 1%. That four-percentage-point shift in managerial perception did not result from training or policy memos. It resulted from structured exposure to evidence over time. Managerial priors are not fixed. They update — but only when the conditions for updating are deliberately created.

The interventions the evidence supports are specific. Equalising developmental feedback across locations is a precondition for fair evaluation, not a supplementary initiative. If the feedback gap is a primary driver of the evaluation gap, closing it must come first — which requires deliberate mechanisms that do not depend on physical proximity to occur naturally. Structured check-ins, documented developmental conversations, feedback equity tracking across locations: none of these are complicated in principle. All of them require organisational commitment that most firms have not yet made.

Calibration processes need to surface location as a variable. Requiring managers to examine whether remote reports are rated differently from in-office reports, and to account for that difference explicitly, introduces friction into a process that currently allows proximity bias to pass through unexamined. Blind review of work products before performance discussions removes one channel through which ambient associations colour formal judgement. Both are achievable without rebuilding the entire performance management architecture.

Output-based criteria, built properly, require the organisational investment in definition and governance that most firms have avoided — not because it is impossible, but because it is slower and more politically demanding than keeping the system already in place. It is also the only intervention that addresses the root cause rather than its symptoms. Defining what good work looks like in each role, at what cadence it should be assessed, and how it should be weighted is a management systems project. It cannot be delegated to an HR initiative or resolved in a workshop. The organisations that treat it as such will produce marginally better versions of the same broken system.

The deeper shift is a redefinition of what presence means professionally. Not physical attendance, but consistent contribution visibility — documented, observable, and equitably distributed regardless of location. Most organisations have built their evaluation cultures entirely around the first. The transition to hybrid work has made the second both necessary and achievable. The gap between those two facts is where the work actually lives.

Get the latest analysis

For decades, organisations ran a substitution so seamless that almost nobody noticed it. Presence stood in for performance — not as deliberate policy, not as a calculated shortcut, but as an assumption so embedded in how workplaces were built that it never needed to be stated. Hybrid work found the joins.

The question it is forcing is not where people should sit. Organisations are still trying to answer that question, with return-to-office mandates and attendance policies that treat location as the variable worth controlling. It is the wrong variable. The right question is what organisations are actually evaluating when they say they are evaluating performance — and whether they are prepared to discover that the answer diverges substantially from what they assumed.

Workers, particularly women and people carrying caregiving responsibilities, are making rational choices inside irrational systems. The cost lands hardest on the people who had the most to gain from flexibility. The organisations that close the gap between fair and accurate will not just run better performance reviews. They will finally be measuring what they always claimed to measure.

That is a more difficult project than it sounds. Most have not yet started.

  1. The Structural Illusion of Control: Organisations have historically traded the difficulty of measuring actual output for the convenience of measuring physical presence, creating a century-long assumption that visibility is a reliable proxy for performance

  2. The Feedback Deficit as a Career Determinant: Professional development is not a static trait, but a cumulative trajectory powered by continuous, informal feedback; when distance severs these inputs, it fundamentally degrades the long-term career potential of the employee

  3. The Rationality of Performative Labour: In any system where accessibility is the primary signal for value, employees will logically prioritise the appearance of availability over the quality of the work, even when such behaviour is actively counter-productive

  4. The Equity Paradox of Flexible Work: Offering flexibility without reforming the underlying evaluation infrastructure does not resolve exclusion; it merely shifts the barrier from the point of entry to the point of advancement

  5. The Legibility Trap of Modern Management: Organisations continue to rely on presence-based evaluation because it provides a clear and low-effort signal, despite the fact that it is a statistically flawed measure that systematically distorts the entire talent pipeline

  6. The Divergence of Fairness and Accuracy: In a hybrid environment, a performance review that feels “fair” to a manager is often objectively inaccurate because it is based on ambient memory rather than documented professional contribution

We have seen that Artificial Intelligence is not coming for the interns—it is coming for the managers. However, as the middle of the corporate ladder thins out, a more unsettling question emerges: If we remove the middle rungs, how does anyone ever reach the top?

In our next article, we shall explore the “Hollowed Pyramid” and the looming crisis of expertise. We will examine what happens when entry-level roles disappear alongside the managers, and why the next generation of leaders might find themselves with nowhere to start.

Stay tuned—the structure of your career is about to change shape.

Stay in the loop

Continue the journey

Thanks for reading The Workforce Lens’s Substack! This post is public so feel free to share it.

Share

Leave a comment

  1. Bloom, Han & Liang (2024), Nature; Emanuel, Harrington & Pallais (NBER Working Paper 31880)

  2. Cullen & Perez-Truglia, American Economic Review; Wharton faculty working paper on remote work and applicant diversity (2023);

  3. Deloitte 2024 Global Human Capital Trends;

  4. Gallup 2024 State of the Global Workplace; Vienna University of Technology / Frontiers in Psychology (2021);

  5. UCLA Breaking Bias / SEEDS model; WFH Research (Bloom, Barrero, Davis)

Read the original on theworkforcelens.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.