~/thinking/turn-rejected-agent-actions-into-a-learning-loop.md
Turn Rejected Agent Actions Into a Learning Loop
Blocking an AI agent's bad action is only half the job.
Meta's expansion of Muse into business workflows got me thinking about the other half. Muse's design splits two roles: the agent proposes an action, and a separate system, Sentinel, decides whether it can use a connector or reach the network.
That boundary is built for safety. From my work on agent governance, I think it's also one of the richest sources of quality data a platform has.
Suppose a business owner asks an agent to email a specific customer segment, but it proposes sending to the entire list. A scope check rejects the action.
Here's the loop I think platforms should make a built-in capability:
→ The agent gets structured feedback on which check failed
→ It revisits its assumptions and narrows the audience to the requested segment
→ The fix is verified against the original task
→ The attempt, the rejection, and the fix become a test case for every future version
Verification is the key step, because the gate isn't always right. An allowed action can still be wrong. A rejection can expose a bad rule, or just reflect the owner's preference.
One caution: feedback should help the agent correct its action within the authorized scope. Every retry must still pass independent authorization checks.
This is where enforcement, observability, and evaluation meet, and where I see much of an agent platform's value.