Most Agentforce and Claude rollouts don’t fail because the technology doesn’t work. They fail because of a predictable set of implementation mistakes — the same handful of issues showing up across different industries, team sizes, and use cases. This guide walks through the challenges that most often arise as teams move from “we’ve decided to do this” to actual configuration, and the best practices that address each one.
If you’ve already made the case internally for Agentforce with Claude as a reasoning model, this is the piece to hand to whoever’s actually building it.
Challenge 1: Treating Model Selection as a One-Time Decision
Teams often configure a model at the org level during setup — Salesforce Default, AWS-Hosted Claude, or Gemini — and never revisit it. But Atlas, Agentforce’s reasoning engine, supports per-agent and even per-subagent model overrides, and different workflows genuinely need different things. A customer-facing service agent handling nuanced, multistep issues benefits from Claude’s depth of reasoning; a simple internal lookup agent may not need that overhead at all.
Best practice: Build a lightweight decision framework — accounting for reasoning complexity, latency sensitivity, and compliance requirements — and apply it per agent as new use cases are added, not just once during the initial rollout. Revisit the framework quarterly as new agents come online, since the right model mix for your org today won’t necessarily be right in six months.
Challenge 2: Building Agents Before the Data Is Ready
This is the single most common root cause behind an agent that technically works in testing but produces unreliable answers in production. Duplicate account records, inconsistent field values, and incomplete Data Cloud profiles don’t just create minor inaccuracies — they undermine the specific promise of an agentic system: that it reasons from real business context rather than guessing.
Best practice: Run a data quality audit scoped specifically to whatever objects and fields the agent will reason over — not a generic org-wide cleanup — before configuration begins. Treat this as a go/no-go gate for moving into build, not a parallel workstream that can catch up later.
Challenge 3: Under-Scoping Governance Until Something Goes Wrong
It’s tempting to treat governance as documentation to produce after an agent is built, rather than a design constraint from the start. This shows up as agents with unclear escalation logic, no defined boundary between actions that need approval and actions that can execute autonomously, and no clear owner once the agent is live.
Best practice: Define escalation paths and action boundaries before configuration, not after — specifically, what the agent can do without approval, what it must hand off to a human, and who’s accountable for reviewing its decisions on an ongoing basis. This should be a signed-off document before go-live, not a retrospective fix after an agent takes an action nobody wanted them to take.
Challenge 4: Misunderstanding Where Data Actually Flows
Claudeforce spans three surfaces — Claude inside Agentforce, Claude inside Slack, and the Salesforce-in-Claude plugin — and they don’t all share the same data flow architecture. Claude via Amazon Bedrock, within the Salesforce Trust Boundary, keeps data within Salesforce’s existing security perimeter; the Salesforce-in-Claude plugin, running on Anthropic’s side of the relationship, doesn’t work the same way. Teams that assume uniform data handling across all three surfaces sometimes discover the gap during a security audit rather than before one.
Best practice: Map data flow explicitly for each surface you plan to use, and get sign-off from security and compliance for each surface — not once for “Claudeforce” as a single blanket approval. This is a slower process upfront and a much faster one when an auditor asks a specific question later.
Challenge 5: Rolling Out Multiple Surfaces Simultaneously
Enthusiasm after a successful demo often leads teams to greenlight Agentforce changes, a Slack rollout, and the Salesforce-in-Claude plugin all at once. Each surface has its own configuration nuances, its own stakeholders, and its own failure modes — running them in parallel multiplies the number of things that can go wrong without multiplying the team’s capacity to catch and fix them.
Best practice: Sequence surfaces deliberately. Pick one contained pilot — a single Agentforce agent or a single sales team on the Claude plugin — get it right, measure it, and use those specific lessons to inform the next surface’s rollout rather than guessing at all of them simultaneously.
Challenge 6: Underinvesting in Change Management
A technically flawless agent that nobody trusts or uses is still a failed implementation. This is often where the actual project effort is most lopsided — teams spend months on configuration and days on adoption, only to be surprised when usage numbers come in at the low end three months after launch.
Best practice: Involve the actual end users — sales reps, service agents, admins — in testing from the earliest usable version, not after the “finished” product is ready. Build a feedback loop that’s fast enough to act on (weekly, not quarterly) during the first months post-launch, and treat early friction reports as implementation bugs to fix, not user error to train around.
Challenge 7: Getting Surprised by Consumption-Based Costs
As pricing shifts toward consumption-based models, teams that don’t actively monitor usage per agent or team can find costs scaling in ways that weren’t anticipated during budgeting — particularly once an agent that works well gets organically adopted well beyond its original pilot scope.
Best practice: Set up usage monitoring and budget alerts per agent or team from day one, not after the first unexpectedly large invoice. Treat cost data as a signal worth reviewing alongside adoption and quality metrics, not a separate finance-only concern.
Challenge 8: No Plan for Ongoing Monitoring After Go-Live
Agent behavior isn’t static. New agents get added, models get updated, and what worked well at launch can drift as the underlying data or business context changes. Teams that treat go-live as the finish line often don’t notice degraded agent performance until a customer or employee complains.
Best practice: Assign clear, ongoing ownership for monitoring agent output quality, reviewing which models are running which agents, and auditing data access — on a defined cadence, not an ad hoc one. This is the same discipline as any production system monitoring, applied to a system that reasons rather than just executes fixed logic.
Building an Implementation Team That Can Handle This
Most of the challenges above aren’t purely technical — they’re sequencing and governance-discipline problems, which means the team executing the implementation needs both Agentforce configuration skills and the organizational discipline to slow down at the right moments (data readiness, governance sign-off) even when there’s pressure to move fast. That combination is harder to find than general Salesforce configuration experience alone, which is why many enterprises bring in a partner with specific Agentforce and Claude implementation experience — such as ABSYZ — rather than assembling that discipline internally for a first attempt at this kind of rollout.
Implementing Agentforce with Claude as your reasoning model? Work through the eight challenges above with your team before configuration starts — nearly every one of them is far cheaper to address in planning than to fix after go-live.
