Chapter 6 walkthrough from The AI Contact Center Handbook by Sho Shimoda. Available on Amazon.
Rachel Nakamura, 8:47 PM Sunday, kitchen table
Rachel opens her laptop, pours a second glass of white wine, and starts building next week's schedule for four hundred contact-center agents across three time zones. She has done this every Sunday for eleven years. The spreadsheet — actually three spreadsheets, linked by a set of VLOOKUP formulas she wrote in 2017 and has been terrified to touch ever since — is the master document that decides when four hundred people will get out of bed on Monday, when they will eat lunch on Wednesday, and whether the single mother in her Phoenix office who asked for Thursday off to see her daughter's school play will get it.
Rachel's official title is Senior Workforce Analyst. On the org chart she reports into Operations. In practice she is the person the vice president of the contact center calls at 6:15 AM on a Tuesday when a storm system in the Southeast has just knocked out power to forty thousand customers and the queue is about to explode.
She takes a sip. The forecast for the coming week, generated on Friday by the workforce-management platform the company bought in 2019 from Aspect — now called Alvaria, because Aspect merged with Noble Systems and rebranded — says call volume will be roughly 47,000 contacts, evenly distributed Monday through Friday with a small spike Wednesday morning. Rachel does not believe the forecast. It is based on a rolling 13-week average and it does not know that marketing is launching a promotional email at 7 AM Monday. Rachel knows. Rachel always knows. But there is no field in the platform for Rachel's gut. She starts fixing the forecast by hand. This will take her about three hours.
Workforce Engagement Management is the operational muscle that decides who works, when, and how well — and it is being quietly rebuilt from Excel and gut instinct into a system that learns.
What WEM is (and why it used to be called WFM)
Workforce Engagement Management — WEM for short — is the family of software products that runs the operational side of a contact center. Scheduling. Forecasting. Quality assurance. Performance coaching. Compensation modeling. Employee-engagement surveys. The set of things you would call plant operations in a manufacturing plant, or clinical operations in a hospital. In a contact center it has this specific compound acronym.
Ten years ago the acronym was different — it was WFM, Workforce Management, and it was mostly two things: forecasting call volume and building schedules. Then, around 2014–2016, the industry noticed that scheduling and forecasting was only about a third of what these platforms actually did. Quality management — recording, sampling, scoring, coaching — was the other big chunk. Performance management, gamification, and employee engagement had been getting quietly bolted on for years. Verint renamed "Impact 360 Workforce Optimization" to "Verint Workforce Engagement." NICE renamed "NICE WFO" to "NICE Workforce Engagement Management." Aspect (now Alvaria) followed. Calabrio, the Minneapolis-based challenger, coined the phrase and pushed it into the analyst community. Gartner formalized the category name in a Magic Quadrant retitled from "Workforce Optimization" to "Workforce Engagement Management" in the 2020 edition.
WFM was software that answered "who works when?" WEM is software that answers "who works when, how well are they doing it, and how engaged are they in the work?" The scope grew from scheduling into the whole employee lifecycle inside the contact center.
| Vendor | Positioning in 2026 | Where you see it |
|---|---|---|
| NICE | Market share leader; WEM bundled inside CXone | Large enterprises, all-in-one buyers |
| Verint | Historic heavyweight, strong QA/analytics | Banking, regulated verticals |
| Calabrio | Fastest-growing cloud-native challenger | Fresh cloud-first deployments |
| Alvaria (Aspect + Noble) | Large installed base, ceding share to cloud natives | Banking, telecom legacy estates (Rachel's employer) |
| Genesys | Native WEM inside Genesys Cloud | Genesys-first stacks |
| Amazon Connect | Partners with Calabrio / NICE rather than building | AWS-first estates |
Rachel's employer runs Alvaria — the old Aspect suite, upgraded to a hosted version but architecturally the same product she started using in 2017. That is one reason her Sunday night looks the way it does. On a modern cloud-native WEM, the work Rachel does by hand would either be automated or would be a much lighter-touch review of what the machine already produced. She is doing manually what the software should be doing itself. That is the shape of the operational modernization opportunity in most contact centers today.
Forecasting — from Erlang C to machine learning
Now zoom in on the technically most interesting piece of Rachel's Sunday night: the forecast. For roughly a century, contact centers have forecasted staffing using a formula called Erlang C. Agner Krarup Erlang derived it around 1917 while working at the Copenhagen Telephone Company, needing to figure out how many circuits the city needed to keep hold times acceptable. Three inputs — expected calls, average handle time, target service level — one output: the minimum number of agents needed to hit the target. Every WFM platform in the world uses some variant of it.
Erlang C is a probability formula that tells you how many people you need on the phones to answer a given number of calls within a given wait-time target. It assumes calls arrive randomly, agents work at an average handle time, and no one gives up and hangs up. Those assumptions are almost right, most of the time, which is why the formula has survived a century.
The formula has three known weaknesses that anyone who has used it professionally will recognize. It assumes calls arrive independently in a Poisson distribution — in practice they cluster (billing statements mail on the first of the month; everyone calls on the second). It does not model abandonment — the variant called Erlang A fixes that, but many older platforms still run pure C. And it is fundamentally single-channel; modern centers handle voice, chat, email, messaging, and video simultaneously, each with different service-level targets and agent-multitasking rules. Modeling a modern multi-channel operation with pure Erlang C is like using a bathroom scale to weigh an airplane.
Machine-learning approaches fix all three problems and add capabilities Erlang C never had. Modern stacks use gradient-boosted trees or LSTM neural networks trained on years of historical data, incorporating hundreds of features: day of week, hour of day, weather, marketing campaigns, product launches, competitor outages, macroeconomic signals. NICE's Enlighten AI and Calabrio's Trueforecast both claim 20–40% forecast-accuracy improvements over Erlang-based methods — which translates into either lower staffing cost (if you tighten the buffer) or better service levels (if you keep it and place staff more precisely).
The subtler impact is that ML forecasting can absorb the kinds of adjustments Rachel makes by hand. When she reasons "the marketing email drops at 7 AM Monday and will add about 1,800 contacts by 10 AM," she is doing an intuitive feature-engineering step — adding a marketing-campaign signal to her forecast. A modern ML platform accepts that signal as an input: she updates a "known events" calendar with the campaign details, and the model incorporates it automatically. Her expertise gets encoded into the system rather than trapped inside her spreadsheet.
The transition is not universally welcomed. Rachel and the workforce analysts like her at every enterprise contact center have spent years developing gut-instinct forecasting skills the industry rewarded. When those skills get absorbed into a model, the analyst's job changes. It becomes more supervisory: validating the model's output, curating the signals that feed it, flagging edge cases the model does not recognize. Some analysts embrace that shift. Some see it as a demotion. Managing that transition is a real change-management challenge, and the vendors who succeed with WEM modernization spend as much time on it as they do on the technology.
Intraday intelligence — from an hour and a half to ninety seconds
Fast-forward to Tuesday morning at 6:45. Rachel's finished schedule from Sunday night is loaded. Agents are logging in from Phoenix, Dallas, and Charlotte. The forecast for Tuesday morning called for a moderate start — about 1,200 contacts between 7 and 10 AM. At 6:52 a category-two hurricane makes landfall in South Carolina. By 7:15 the Charlotte property-insurance queue has forty-seven callers and a nine-minute wait. By 7:30, sixty-eight callers, eleven minutes.
This is the moment intraday management earns its salary. Where do the extra agents come from, and what will we not be doing while they are on the storm queue? In the legacy world this looks like phone calls and emails. Rachel messages the auto-insurance supervisor, the general-service supervisor, the escalations supervisor. She negotiates. She sends an all-hands email offering voluntary overtime. She calls the Manila outsourcer. The whole process takes an hour and a half, and by the time reinforcements arrive the queue has been critical for two hours.
Intraday intelligence, done right, compresses that hour and a half to about ninety seconds. The moment queue depth crosses a defined threshold, the platform generates a set of proposed reallocations: shift twenty agents from auto to property, activate seven agents on standby, deploy an SMS to sixty cross-trained agents in home offices offering a two-hour extension shift at overtime rate. Each proposal comes with an impact estimate — "shifting twenty agents from auto raises average auto wait to 3.4 minutes, still within SLA" — and Rachel approves, adjusts, or overrides with a click.
| Intraday metric | Legacy WFM | Modern intraday module |
|---|---|---|
| Time-to-recover from queue event | ~47 minutes (industry avg.) | ~8 minutes (large U.S. health insurer, Calabrio, 2022) |
| Service level on volume spikes | 62% | 88% |
| Overtime cost | Baseline | -23% (smarter about which agents to ask first) |
| Analyst role during a spike | Firefighter — running the decision loop in her head | Fire chief — supervising the system's decisions |
Intraday intelligence does not replace the workforce manager. It moves her from executing decisions to designing the decision framework.
The transition takes work. The intraday module is only as good as its inputs: agent skill data has to be accurate and current, break rules have to be codified, and union or works-council agreements that govern schedule changes have to be encoded properly. Centers that skip that upfront work end up with a system that proposes moves the operations team cannot legally make, which quickly erodes trust in the tool. But centers that do the upfront work correctly find the ROI is fast — usually less than a year — and durable.
Automated QA — from 2% sampling to 100% monitoring
In the traditional model, a QA team samples a small percentage of interactions — typically 1–3%, occasionally up to 5% in high-regulation environments — listens or reviews by hand, and scores against a rubric of fifteen to thirty items. At a 500-seat center handling 300,000 contacts a month, 2% sampling means QA reviews about 6,000 interactions — roughly one review per agent per week if the team is well-staffed. Most are not, and the real number is closer to one review every two to three weeks.
The problem is not the ratio. It is the statistics. Sampling 2% of interactions to score an agent is like assessing a baseball player on two at-bats out of a hundred. The sample size is too small to be meaningful for individual agents, and the delay between the coached call and the coaching feedback — typically four to six weeks — is too long for the feedback to be useful.
Automated QA monitors 100% of interactions. Three-component stack: high-quality speech-to-text transcription (Whisper and its successors from OpenAI, plus comparable models from Google, Deepgram, AssemblyAI — word error rates on English contact-center audio down from around 20% five years ago to under 5% today for the best commercial models); natural-language-understanding models that extract intent, sentiment, topics; and scoring engines that apply the QA rubric consistently across every interaction. Commercial products: NICE Enlighten Autoscoring, Verint's AI Automated Quality Management, Calabrio Analytics with Quality Automation, Observe.AI, Cresta, Level AI. The best of them can score against a full rubric in near real time — the score is available within minutes of the call ending, at accuracies within a few percentage points of experienced human scorers on most rubric items.
There is a critical nuance. Automated QA works best on objective rubric items — was the greeting used, was the disclosure said, was the verification completed, was the correct product mentioned. It works less well on subjective ones — was the agent empathetic, did they handle frustration well, was the interaction warm. The best implementations use automation for the objective items and reserve human scoring for the subjective ones, at a hundred times the coverage on the objective side.
The technology has a failure mode vendors sometimes gloss over. The scoring model is trained on historical human scores, which means it inherits whatever biases those scores contained. If the previous QA team was systematically stricter on non-native English speakers, or on agents in a particular geography, the automated system will learn that pattern and reproduce it. Best-practice deployments include ongoing bias audits, calibration sessions between the automated system and a diverse group of human reviewers, and explicit tracking of score distributions across agent demographic segments. This is not optional. It is what separates a mature deployment from a legal exposure.
How agents actually feel about being scored by AI
We have talked about the technology, the economics, and the operations. We have not yet talked about what it feels like to be Priya, three years into her contact-center career, when her employer announces every one of her calls is now going to be scored by an AI. The answer is: complicated.
The initial reaction, in every deployment I have observed or read case studies of, is anxiety. The old QA process was uncomfortable but knowable — a call might be reviewed, most would not be. The felt experience of the job included a certain amount of privacy. When 100% of calls become subject to scoring, that felt privacy disappears. Agents report a sense of being watched constantly, and for the first few weeks average handle time actually goes up as agents self-consciously verify every rubric item. That anxiety attenuates over two to three months. Whether it fully disappears depends on how the organization implements the tooling.
| Pattern | What agents describe | Verdict |
|---|---|---|
| Transparent rubric, live score | Feedback — they can see the criteria and their own scores in near-real time | Tolerable |
| Coaching as the primary use | Useful — agents bring specific interactions to their supervisor | Tolerable |
| Aggregate reporting only | Trusted — individual scores private, aggregate anonymized data shared | Tolerable |
| Opaque scoring | Accusation — every flag treated as an attack; trust collapses | Demoralizing |
| Auto-triggered PIPs | Fear — low score triggers performance-improvement plan without human review | Demoralizing |
| AI-only appeals | Kafkaesque — challenging a score means talking to another AI | Demoralizing |
The technology of 100% QA is neutral. It can be deployed in a way that treats agents as adults, gives them useful feedback, and improves their working lives. It can also be deployed in a way that treats them as suspects, monitors them into compliance, and grinds down whatever remains of the humanity in the work. The vendors ship the same platform to both kinds of organizations.
The organizations whose agents describe the system positively — and there are plenty of them — share a few characteristics. Their leaders talked with agents before the deployment, not just to them. Launch communications acknowledged the anxiety and named it explicitly. The first three months focused on calibration and refinement rather than performance action. And they told their agents from the beginning exactly what the AI was and was not going to be used for — and then they held to that. The organizations whose deployments went badly did the opposite. They rolled the tool out with a compliance-first message. They started acting on scores immediately. They kept the rubric weightings secret. They used the automation to justify staff reductions before the remaining agents had a chance to understand the new system. The result was predictable — complaints, union grievances, and a spike in turnover that undid the operational savings.
Priya — the six-months-in agent from Chapter 4 who we will keep meeting throughout this book — works at one of the good deployments. She checks her scores each afternoon before her end-of-shift call with her supervisor. She has learned that the AI flags her on account-closure disclosures about 12% of the time, which she now knows is because she says the disclosure a few beats too fast for the compliance model's speech recognizer. Her supervisor helped her fix it. Her scores are up. She has a better relationship with the tool than she had with the old QA process, which used to feel arbitrary. The new system feels, to her, more fair. That "more fair" outcome is the design goal. Whether a specific WEM deployment reaches it depends less on the vendor and more on the organizational choices about how to use what the vendor ships.
What to take with you from Chapter 6
Workforce Engagement Management is the operational layer that runs the contact center's people function — scheduling, forecasting, quality, coaching, engagement. It was called Workforce Management until the vendors realized they were selling something bigger, and the WEM rebrand around 2016–2020 signaled a shift from "who works when" to "who works when, how well, and how engaged."
Two of the biggest technical shifts inside WEM are happening in forecasting and quality assurance. Forecasting is moving from the century-old Erlang C formula to machine-learning models that incorporate hundreds of features, absorb the tribal knowledge of expert analysts as signals, and deliver 20–40% accuracy improvements. That accuracy shows up as either lower staffing cost or better service levels. It also changes what the analyst's job is — from executing the forecast to supervising the model that produces it.
Quality assurance is moving from a 1–3% sample to 100% monitoring, powered by speech-to-text that has become genuinely reliable and by NLU models that can score against a rubric in real time. The economics are dramatic — a large center can reduce QA headcount by 60% or more while increasing coaching coverage by an order of magnitude. But the human implications are subtle. 100% monitoring can be deployed as a coaching tool agents describe as helpful, or as a surveillance tool agents describe as demoralizing. The choice is not the vendor's. It is the organization's.
The through-line across every WEM shift is the same: the systems are moving from batch to real-time, from rules to learning, and from operational execution to operational supervision. The workforce analyst, the QA reviewer, and the intraday manager are not being replaced. Their work is being amplified, and the work that remains is the higher-judgment work the machines cannot yet do — designing the rubric, curating the signals, appealing the edge cases, coaching the humans.
Chapter 7 turns from the operational side to the customer side — hyper-personalization and omnichannel journeys, the muscle that decides how a customer's experience flows across the many surfaces where the brand touches them. If WEM is the story of the people who serve customers, hyper-personalization is the story of the customers being served.