All insights

    AI

    Why Most Internal AI Chatbots Go Unused Within 90 Days

    Big Sky Consulting Group · August 26, 2026 · 7 min read

    The launch went well, which is the part that misleads you

    Week one, usage is strong. Half the company tries it, because a new internal tool is a novelty and the announcement had a demo video. Week three, the number is a third of that. By day 90 you have a dashboard with a flat line near the bottom, a handful of loyal users who mostly work in one department, and a renewal conversation coming in eight months.

    Nobody rejected the tool. There was no revolt, no memo, no meeting where someone said this is bad. People simply stopped opening it, which is the quietest failure mode a software purchase has and the hardest one to raise with the board.

    The reflex explanation is that the model was not good enough, or that the company is not ready, or that change management was underfunded. Occasionally one of those is true. Usually the tool is sitting in the wrong place relative to the work, and it was sitting in the wrong place on the day it shipped. Adoption did not decay. It was never structurally possible.

    Placement beats intelligence, and it is not close

    Here is the pattern, and it repeats across companies that share almost nothing else.

    A person is doing their actual job inside a system. A claims adjuster is in the claims platform. A dispatcher is in the TMS. A finance analyst is in the ERP, or realistically in a spreadsheet exported from the ERP. That system is where their hands are, where their context lives, and where their manager measures them.

    The AI assistant is not in that system. It is a separate tab, or a Teams app, or an icon on an intranet page. Using it costs a context switch: leave the work, restate the situation the other system already knows, read an answer, carry the answer back, retype it. That is fifteen to forty seconds of overhead on top of whatever the model does well.

    So the assistant only wins when the question is expensive enough to justify the switch. Big, rare, genuinely hard questions clear that bar. But most of what people need during a working day is small and frequent, which is precisely the volume that would have made the tool worth buying. The small questions never migrate. What is left is a low-frequency tool that a few people use for a few hard things, which is a real but very modest outcome priced like a transformation.

    This is why the honest predictor of 90-day usage is not model quality or training budget. It is a single question we ask before anyone signs anything: what does the user have to stop doing in order to use this? If the answer involves leaving the screen where the work happens, plan for the flat line.

    The second predictor is what the tool knows. An assistant that can answer from the policy library but cannot see this customer, this shipment, or this claim is answering a general question in a specific moment. People forgive that once. They do not build a habit around it.

    The comparison you are actually losing to

    Your internal assistant is not competing with the old intranet search. It is competing with the phone in your employee's hand.

    PagerDuty's 2026 workplace survey, run by Wakefield Research across 1,250 office professionals at companies above $500 million in revenue, found that 66% had used AI tools at work they believed were not permitted, and that 89% first encountered the AI tools they use for work in their personal life. Gartner's figure for unauthorized AI use at work moved from 41% in 2023 to 68%.

    Read those numbers as a product review rather than a security incident. Your people are not avoiding AI. They are choosing a different one, fluently, on their own time, and the one they chose has no procurement process and no onboarding session. When your sanctioned tool loses that comparison, it loses to something free that they already know how to prompt.

    That reframes the question. The problem is rarely getting employees to use AI. It is that the version you paid for has to be better at their job than the version they already like, and "better at their job" almost always means better placed and better informed rather than better reasoning.

    This is the general shape of the problem. Which parts apply to your process depends on answers only your systems can give.

    Put us on it, from $5,000

    Why the pilot did not warn you

    Pilots are run by people who volunteered. That population is enthusiastic, tolerant of friction, and often selected because they were already using AI on their own. They will absorb a context switch without noticing, and they will write generous feedback.

    The general population will not. So the pilot measures the ceiling and the rollout measures the floor, and the gap between them shows up as a usage cliff somewhere around week three.

    MIT's State of AI in Business 2025 study, based on 52 executive interviews, 153 surveyed leaders, and 300 public deployments, found that roughly 95% of generative AI pilots produced no measurable P&L impact, and attributed it primarily to tools that did not retain context or improve inside the workflow rather than to weak models. That finding gets quoted as a scare statistic. We read it as a placement statistic. Something that cannot hold the context of the work will not become part of the work.

    This is the same failure family we described in when AI is the wrong answer to an operations problem: the technology functions exactly as advertised, and the surrounding process was never arranged to receive it. It is also the reason we push buyers to ask vendor evaluation questions about integration depth and permissions before questions about model capability. A demo answers questions in a sandbox. Your users will ask questions in a queue, under a handle time target, with a customer waiting.

    What the ones that survive have in common

    We do see these work. The ones that hold usage past a year tend to share three unglamorous traits.

    They are narrow. Not "ask anything about the company." One job, for one role, defined tightly enough that the tool is right nearly every time in its lane. Narrow tools earn trust faster because their failure boundary is legible. Users learn what it is for in a week instead of guessing for a quarter.

    They are embedded where the work already happens, not adjacent to it. Inside the ticket, next to the claim, in the field the person was going to fill in anyway. The best ones are barely experienced as chat at all.

    They are wired to live data with real permissions, so the answer reflects this record rather than the general policy. This is the expensive part, it is the part vendors quote thinly, and it is the part that decides the outcome. It is also where the honest version of the project stops being a chatbot purchase and starts being an integration project with a chat interface on top, which is a different budget line and a different conversation with your CFO. We walked through how those omitted line items move the number in our piece on recalculating automation ROI.

    Notice what is not on that list. Nothing about model selection. Nothing about prompt training for staff. Nothing about an internal champions program. Those help at the margin, and they cannot rescue a tool that sits in the wrong place. You can lead a team to a chatbot, but you cannot make them think it is worth the extra tab.

    The uncomfortable version of the advice

    Sometimes the right call at day 90 is to stop, not to relaunch. A tool with 40 committed users in one function and none anywhere else is not a failed rollout, it is a correctly sized departmental tool wearing an enterprise price tag. Renegotiate to what it actually is. That is usually a better trade than a second launch campaign aimed at people who already made their decision quietly.

    And sometimes the right call before day zero is not to buy at all. If your knowledge base is stale, ownerless, and contradicts itself in three places, an assistant will retrieve stale, ownerless, contradictory answers with excellent grammar. It will lose the trust of your staff in about two weeks, and after that no amount of cleanup buys it back. Fix the source material first, on its own timeline, and evaluate the assistant afterward against a corpus worth searching.

    Which of these applies to you depends on things an article cannot know: where your people actually spend their day, what your systems will let a tool read and write, whether the questions being asked are frequent and small or rare and large, and what your current contract lets you renegotiate. Those four answers decide it, and they take about a week to gather properly.

    If you are staring at a flat usage chart, or you are three weeks from signing for a tool that will produce one, that is worth a conversation before the renewal date decides it for you. Book a consult and we will look at where the tool sits relative to the work, what it can see, and whether the honest fix is integration, narrowing, renegotiation, or walking away.

    AI AdoptionChange ManagementBuy-Side AdvisoryInternal Tools

    Want this applied to your business?

    A paid working session on one of your processes, ending in a written implementation plan. Not a sales call. Consults start at $5,000, and you see the price before you commit.

    Book a consult