The Claim on the Table
A recent Forbes piece introduces what the author terms RHLCF - Reinforcement Human Learning from Computer Feedback - framing it as the natural successor to RLHF (Reinforcement Learning from Human Feedback). The argument is structurally straightforward: just as humans once provided feedback signals to train AI systems, AI systems will increasingly provide feedback signals to train humans. The piece positions this as evolutionary, even optimistic. I want to take that claim seriously, because the organizational theory implications are considerably less tidy than the framing suggests.
What Feedback Authority Actually Does in Organizations
Classical organizational theory treats feedback as a control mechanism that flows from principals to agents. Managers evaluate workers; workers adjust behavior. The authority embedded in that feedback loop is not incidental - it is the mechanism through which organizations socialize competence and enforce norms. When you invert the feedback direction, as the RHLCF framing proposes, you are not simply changing the communication channel. You are relocating the source of evaluative authority from human judgment to algorithmic output. Rahman (2021) describes a version of this dynamic in the platform economy as the "invisible cage" - workers who nominally retain autonomy but whose behavioral parameters are continuously narrowed by algorithmic feedback structures. RHLCF, if adopted at scale, does not limit this dynamic to gig workers. It extends it into conventional employment.
The Awareness-Capability Gap Gets Structurally Embedded
The RHLCF framing assumes that receiving AI-generated feedback will improve human performance. This assumption deserves scrutiny. Research on algorithmic literacy consistently finds what I have called the awareness-capability gap: knowing that an algorithm is shaping your environment does not translate into knowing how to respond effectively to it (Kellogg, Valentine, and Christin, 2020). The RHLCF model implicitly assumes workers will develop accurate schemas of what the AI system is optimizing for, then calibrate their behavior accordingly. But Gagrain, Naab, and Grub (2024) find that most workers develop folk theories rather than accurate structural schemas - plausible but often incorrect models of how algorithmic systems operate. If workers are receiving feedback from systems they fundamentally misunderstand, the feedback loop does not produce genuine competence development. It produces behavioral compliance shaped by misread signals.
Routine Versus Adaptive Expertise Under Inverted Feedback
Hatano and Inagaki (1986) distinguish between routine expertise - the ability to execute procedures reliably within a stable environment - and adaptive expertise - the ability to recognize when structural conditions have changed and adjust accordingly. The RHLCF model, as described in the Forbes piece, is optimized for producing routine expertise. AI feedback systems reward consistency with past patterns. They are not well-suited to signaling when a worker should deviate from those patterns because the underlying environment has shifted. This is not a minor limitation. It is a structural one. Organizations that build feedback architectures around AI evaluation may find they are systematically developing workers who are highly responsive to current algorithmic criteria and poorly equipped to transfer that competence when platform logic changes. The variance puzzle I track in my dissertation research - why workers with identical access produce dramatically different outcomes - is not solved by giving everyone the same AI feedback. It may in fact be amplified, as initial behavioral differences get locked in through differential reinforcement.
The Governance Question the Piece Does Not Ask
The Forbes framing treats RHLCF as a workplace evolution story. The governance story is harder. Hancock, Naaman, and Levy (2020) note that AI-mediated communication systematically alters the perception of message authenticity and authority. When feedback originates from an AI system rather than a human supervisor, workers may assign it different epistemic weight - either dismissing it as mechanical or, more problematically, over-crediting it as objective. Neither response is accurate, because AI feedback systems encode the values and priorities of whoever designed them. Sundar (2020) frames this as machine agency - the tendency for users to attribute autonomous judgment to systems that are, in fact, expressing human-designed optimization targets. RHLCF does not eliminate the politics of evaluation. It obscures them behind an interface that signals neutrality it cannot actually deliver.
What This Means Practically
The question worth asking about RHLCF is not whether AI can provide useful performance feedback - it clearly can in bounded domains. The question is what organizational theory predicts happens when feedback authority migrates to systems that workers do not structurally understand, that optimize for historical patterns rather than adaptive capacity, and that represent human evaluative priorities as objective machine outputs. The answer, drawing on Schor et al. (2020) and Rahman (2021), is a form of platform-style dependence extended into the interior of conventional employment. That may or may not be the future of work. But it should be analyzed as a governance structure, not celebrated as an evolutionary step.
Roger Hunt