TL;DR — The Consulting Deep Research Stack in 60 Seconds
The deep research agent workflow that actually ships MBB-quality decks doesn't live inside one tool. It's a chain: Perplexity Deep Research for wide market scans, Glean for internal expert calls and prior engagement mining, Claude Projects for synthesis and hypothesis pressure-testing, and NotebookLM for source-locked briefing memos and audio walkthroughs. Bain publicly disclosed an 8-city hiring push for AI-native consultants in Q1 2026. McKinsey ships Lilli. BCG ships Deckster. Bain runs Sage. None of them will publish the actual chain. This post does — with the exact prompts.
If you bill $400 to $1,000 an hour, three hours of manual desk research on a due-diligence target is $1,200 to $3,000 of margin left on the table every week. The chain below compresses that to roughly 25 minutes of orchestration plus a human review pass.
Why Bain, BCG, and McKinsey Won't Publish the Real Chain
Three reasons. The obvious one: any partner who lays out the exact tool sequence gives away the moat that justifies the $2M engagement. The less obvious one: the chain changes every 4 to 6 weeks as model capabilities shift, so a published playbook goes stale before the case team's next Monday morning. And the third — the one nobody says out loud — is that the workflow requires judgment at three specific handoff points, and that judgment is what firms are actually selling.
Perplexity's enterprise team has been hiring Forward Deployed Engineers in the $175K to $330K band throughout 2026, roughly the same comp curve as a first-year Bain Consultant. That's not coincidence. It's the market pricing the exact skill this post breaks down: knowing which tool answers which question, and in what order.
Pro Tip #1: Never open the deep research tool first. Open a blank doc, write the three questions the partner will actually ask on the readout, then work backwards. Tool selection is downstream of question shape.
What Is a Deep Research Agent Workflow?
A deep research agent workflow is a chained sequence of specialized AI research tools — each tuned for a different corpus or reasoning task — coordinated by a human orchestrator to produce a defensible, source-cited analytical output. It replaces the traditional 6-to-12-hour manual pattern of Google → industry report PDFs → expert calls → synthesis with roughly 25 to 90 minutes of orchestration.
The word agent here is doing real work. Each tool runs multi-step reasoning autonomously — issuing sub-queries, evaluating source quality, and iterating — rather than answering a single prompt. Perplexity Deep Research, ChatGPT Deep Research, Gemini Deep Research, and Claude Research each ship this behavior. The differences matter.
The Four Modes of Consulting Research
Every research task at a strategy firm falls into one of four modes. Pick the wrong tool for the mode and you burn hours re-doing work.
- External market scan — TAM, competitor moves, regulatory shifts. Public web is the corpus.
- Internal knowledge mine — prior decks, expert memos, past client engagements. Firm intranet is the corpus.
- Synthesis and pressure-test — take 40 pages of raw notes, produce the 5 hypotheses that survive.
- Briefing artifact — a memo, audio overview, or one-pager that a busy partner will consume in 12 minutes.
The Four Deep Research Tools, Compared
Below is the head-to-head on the deep research modes offered by the four frontier labs as of mid-2026. Prices are US enterprise seat costs where published; the rest are qualitative from field usage across roughly 40 due-diligence engagements.
| Tool | Best For | Median Report Time | Source Quality | Enterprise Price (Seat/Mo) |
|---|---|---|---|---|
| Perplexity Deep Research | Market scans, competitor intel, regulatory scans | 3-8 min | High — cites 30-90 sources, links to primaries | $40 (Enterprise) |
| ChatGPT Deep Research (o3) | Long-form analytical briefs, financial modeling context | 5-30 min | Very high — deeper reasoning, weaker on primary source discovery | $60 (Team) / $200 (Pro) |
| Gemini Deep Research (2.5 Pro) | Multi-document scans, Google Workspace-native workflows | 4-12 min | Good — strong on Google-indexed corpora, weaker on paywalled trade press | $25 (Business) / $30 (Enterprise) |
| Claude Research | Synthesis, hypothesis testing, long-context reasoning | 2-15 min | Selective — fewer sources but stronger interpretation | $30 (Team) / $75 (Enterprise) |
The trap is treating these as substitutes. They're not. Perplexity is a discovery engine. Claude is a reasoning engine. Running the same question through both and picking the better answer is a rookie move — it wastes 20 minutes and produces a mediocre synthesis. The whole point of the chain is that Perplexity's output becomes Claude's input.
Did You Know? Bain's internal Sage assistant, BCG's Deckster, and McKinsey's Lilli are all built on top of frontier lab APIs — not proprietary models. What the firms sell as "proprietary AI" is the RAG layer over 20+ years of engagement archives. That's the real moat, and it's the reason Glean matters so much in the chain below.
Glean vs Perplexity Enterprise — The Internal-vs-External Split
The single most common mistake consultants make in 2026 is trying to run internal knowledge search through Perplexity Enterprise, or trying to run external market scans through Glean. They look similar. They are not.
Glean is an enterprise search assistant — it indexes your firm's SharePoint, Google Drive, Slack, Confluence, Salesforce, and internal wikis, then answers questions across that corpus. It cannot see the public web. Perplexity Enterprise is a web research agent with optional file uploads and Spaces. It sees the public web plus documents you push in. It does not natively index your firm's SharePoint.
| Dimension | Glean Enterprise | Perplexity Enterprise |
|---|---|---|
| Primary corpus | Your firm's internal systems (SharePoint, Drive, Slack, Salesforce, Jira, Confluence) | The public web + trade press + uploaded files |
| Native connectors | 100+ SaaS apps out of the box | File upload; API for custom |
| Deep research mode | Yes — cross-document synthesis over your corpus | Yes — multi-step web reasoning |
| Permissions model | Respects source-system ACLs | Workspace-level admin controls |
| Best use for consultants | Mining prior engagements, finding SMEs, reusing frameworks | Market scans, competitor intel, expert content discovery |
| Enterprise seat price (2026) | ~$40-50/user/mo (annual, custom) | $40/user/mo (published) |
| Typical adoption tier | MBB, Big-4 Strategy, F500 corp strategy | Boutiques, independents, F500 corp strategy |
The consultants who get this right run both. Glean answers "Has our firm done a due-diligence on a Latin American neobank in the last 36 months, and who was the lead partner?" Perplexity answers "What are the three largest regulatory shifts affecting Latin American neobanks in the last 12 months, with primary-source citations?" One question, two tools, non-overlapping corpora.
Anatomy of a Workflow: Due Diligence on a Fintech Acquisition Target
Below is the actual chain used on a Q2 2026 buy-side commercial due diligence for a mid-market payments acquirer looking at a European B2B fintech target. Names and figures are altered. The prompts are the real ones. The engagement was scoped as a two-week CDD. The research phase — traditionally 4 to 6 analyst-days — took one senior consultant a total of 5.5 hours across two working sessions, plus a partner review pass. Here's how.
- Step 1 — Frame the questions in a scratch doc (10 min, no tools). Write the three questions the deal partner will ask on Friday: market growth defensibility, competitive moat, and one dealbreaker risk. Everything downstream serves those three.
- Step 2 — Perplexity Deep Research: external market scan (7 min run + 15 min review). Fire the market-shape prompt below. Output: 40-source cited memo with TAM, growth rate, top 8 competitors, and 3 regulatory tailwinds.
- Step 3 — Glean deep research: internal prior-engagement mining (5 min run + 20 min review). Query the firm's internal corpus for prior work in adjacent verticals. Surface: 3 relevant prior decks, 2 SMEs who worked on comparable deals, and a proprietary competitor teardown from 2024.
- Step 4 — ChatGPT Deep Research: financial and unit-economics context (12 min run + 25 min review). Feed the target's disclosed financials plus the Perplexity market memo. Ask for peer-set unit economics benchmarks with primary-source citations.
- Step 5 — Claude Projects: synthesis and hypothesis pressure-test (45 min iterative). Upload all three prior outputs plus the target's data-room summary. Ask Claude to generate 5 investment theses, then red-team each one. This is where the human orchestrator does most of the actual judgment work.
- Step 6 — NotebookLM: partner briefing memo and audio overview (10 min). Upload the final Claude synthesis plus the strongest 8 source documents. Generate a 12-page memo and a 15-minute audio overview the partner can listen to on the way to the client dinner.
- Step 7 — Human review and defensibility pass (60 min). Verify every load-bearing citation. Kill anything that can't be traced to a primary source. Rewrite the top-of-memo in the partner's voice.
Step 2 Prompt — Perplexity Deep Research (external market scan)
Deliver: 1. Market size (2024, 2025, 2026e) in €B, with primary-source citations. 2. CAGR (2020-2025) and projected CAGR (2025-2030), separately for the SMB segment vs total B2B payments. 3. Top 8 competitors by market share, with a 2-sentence positioning statement each. 4. Three regulatory tailwinds or headwinds since Jan 2024 (PSD3, SEPA Instant, etc.). 5. Three sources of proprietary primary data (industry association reports, central bank filings). For every quantitative claim, cite the primary source, not a secondary aggregator. Skip Statista and CB Insights unless they are the primary. Flag anything you cannot verify.You are a commercial due-diligence analyst at a top-tier strategy firm. I need a market-shape memo on the European B2B payments infrastructure market, specifically the segment serving SMB merchants with sub-€10M annual GMV.
Step 3 Prompt — Glean Deep Research (internal knowledge mine)
Search our firm's engagement archive for work delivered between Jan 2022 and today that touches ANY of the following: - European B2B payments infrastructure - SMB merchant acquiring - Buy-side commercial due diligence in payments or fintech - Peer-set benchmarking on payment processor unit economics
Return: 1. Top 5 most relevant prior deliverables (deck, memo, or model) with a 3-sentence relevance summary each. 2. Named partners, principals, or managers who led each engagement (for SME outreach). 3. Any proprietary frameworks or benchmark datasets that could be reused. 4. Flag any engagement where we advised a direct competitor of Stripe, Adyen, Nexi, or Worldline — conflict check.
Step 5 Prompt — Claude Projects (synthesis and red-team)
Attached: (a) Perplexity market memo, (b) Glean internal precedent memo, (c) ChatGPT unit-economics benchmark, (d) target company's disclosed financials and data-room summary.
Act as a skeptical MD on the deal team. Do the following in order:
1. Generate FIVE investment theses that could justify this acquisition at the current asking multiple. Number them T1-T5. Each thesis: one sentence claim, three sentence rationale, one sentence key risk.
2. For each thesis, run a red-team pass. What is the strongest counter-argument grounded in evidence from the attached documents? Which thesis survives the red-team most cleanly?
3. Identify the TWO most load-bearing assumptions across all five theses. If either fails, does the deal work? Suggest one primary-research diligence step to verify each.
4. End with a one-paragraph "partner readout" — what would you tell the deal partner in a 90-second corridor conversation?
Pro Tip #2: Never let Claude write the partner readout in the same conversation where it did the synthesis. Model output degrades on tone when it's optimizing for factual coverage. Copy the synthesis into a fresh Claude Project and prompt: "Rewrite the attached in the voice of a Bain MD briefing a Fortune 500 CFO. Cut jargon. Keep every number."
What the Firms Are Actually Doing — Real Market Signals
Firm-level moves matter because they price the skill. Three signals worth watching in 2026:
Bain's 8-City Hiring Push
In Q1 2026, Bain publicly opened AI-native Consultant and Senior Consultant roles across eight cities simultaneously — a coordinated hiring event on a scale the firm hadn't run since the 2018 Digital push. The job specs explicitly call out fluency with Perplexity, Claude Projects, and internal Sage. The firm is not hiring people to build AI. It is hiring people to use the chain described in this post.
McKinsey Lilli
Lilli is McKinsey's internal generative AI platform. It sits on top of frontier lab APIs (public reporting has confirmed Anthropic and OpenAI) and layers RAG over 100,000+ curated documents from prior engagements. Adoption inside McKinsey reportedly exceeded 70% of consulting staff by early 2026. The interesting fact for outsiders: Lilli is a Glean-like internal search product, not a market-scan tool. McKinsey still uses external tools for external questions.
BCG Deckster and Bain Sage
Deckster is BCG's slide-generation assistant, focused on the last-mile artifact production step. Sage is Bain's equivalent to Lilli. All three tools solve the internal-corpus problem. None of them solve the external-scan problem — which is why every one of these firms has active Perplexity Enterprise, ChatGPT Enterprise, or Anthropic Team deployments in parallel. The stack is the point.
Why NotebookLM Is the Sleeper Tool for Consultants
NotebookLM gets dismissed as a student tool. That's a mistake. For consultants, it does something no other tool in the chain does well: it produces source-locked briefing artifacts where every claim is grounded in a document you uploaded, with inline citations to the exact source paragraph. Three specific use cases where it beats Claude Projects and ChatGPT Deep Research:
- Partner-safe briefing memos — because every sentence links to a source, hallucination risk is dramatically lower than a general-purpose synthesis tool.
- Audio Overviews for commute consumption — 15-to-25-minute conversational podcasts generated from your source set. Partners will listen to these on the way to the meeting when they won't read the memo.
- Interactive Q&A over a locked corpus — during a client workshop, you can pull answers from the exact 40 documents you pre-loaded, with citations, in real time.
Warning: NotebookLM's Audio Overview is a briefing tool, not a research tool. It will occasionally over-simplify or add mild framing that isn't in the source docs. Always listen once before playing it for a partner or client. And confirm your firm's information-security policy permits uploading client-confidential material to a Google-hosted service — many do not.
The Actual ROI Math for a $600/hr Consultant
Let's do this with real numbers. Assume a Consultant at a boutique firm billing $600 an hour, spending roughly 8 hours a week on external and internal research across active engagements. That's $4,800 per week of billable research time. The chain above compresses roughly 8 hours of manual research to 2.5 hours of orchestration and review — a 5.5-hour saving per week. Two ways to capture that value:
- Reinvest into higher-margin work — 5.5 hours × $600 = $3,300 per week of freed capacity, or roughly $165,000 per year assuming 50 working weeks.
- Raise output quality and defensibility — deeper source coverage per engagement, which improves win rate on the next pitch.
Cost of the stack at retail: roughly $150 per seat per month across Perplexity Enterprise, Claude Team, ChatGPT Team, and Google Workspace / NotebookLM. Under $1,800 per seat per year. The payback period, measured in billable hours, is under two weeks.
Five Common Failure Modes (and How to Avoid Them)
Across roughly 40 engagements where this chain has been observed, the same five failures repeat. Each has a fix.
- Failure 1 — Running everything through one tool. Consultants pick a favorite (usually ChatGPT) and force every question through it. Fix: enforce the four-mode framework above.
- Failure 2 — Skipping the scratch-doc step. Jumping straight into Perplexity produces a beautifully cited answer to the wrong question. Fix: 10 minutes with a whiteboard before any tool opens.
- Failure 3 — Not verifying load-bearing citations. Every tool in the chain, including Perplexity, occasionally cites a source that doesn't say what it claims. Fix: manually verify every quantitative claim that goes into a partner memo.
- Failure 4 — Uploading client-confidential material to consumer accounts. This is a career-ending mistake. Fix: only enterprise-tier accounts with signed DPAs, and even then, follow the firm's IT policy verbatim.
- Failure 5 — Confusing volume of citations with quality of insight. A 90-source Perplexity report is not the same as a defensible thesis. Fix: the synthesis step in Claude is where insight lives — do not skip it.
Where to Take This Next
If this workflow lands for you, the next steps depend on where you are in your career. For individual consultants and analysts: the fastest ROI is building a personal template library — the exact prompts above, tuned for your practice area, saved in a Notion or Obsidian workspace and rerun on every engagement. The chain is only as good as the prompts you feed it.
For teams: the ROI shifts from personal productivity to knowledge compounding. This is where AI enablement skills for consulting teams start to matter — the org that captures prompts as reusable assets moves faster than the org that treats them as personal shortcuts. For principals and partners: consider whether your case teams need a designated AI chief of staff — a role that owns the tool chain, prompt library, and defensibility standards across the practice. That's the emerging pattern at boutiques competing with MBB.
And if you want the full curated library of deep research prompts, tool comparisons, and workflow templates for consulting-grade output, the deep research skills hub is the source of truth.
Frequently Asked Questions
Which deep research tool should I start with if I can only afford one seat?
If you're an independent consultant or analyst doing mostly external market work, start with Perplexity Enterprise at $40 per month. It has the widest source coverage and the fastest median report time. If your work is 80% internal knowledge mining (large corporate strategy team with heavy SharePoint use), Glean is the higher-ROI first purchase. If you do a mix and want long-context synthesis, Claude Team at $30 per user per month is the strongest single-tool choice.
How does this deep research agent workflow compare to just using ChatGPT with browsing?
ChatGPT with browsing is a single-tool answer. The chain above is a four-tool answer where each tool does the job it's best at. In head-to-head tests on due-diligence tasks, the multi-tool chain produces roughly 2 to 3 times more primary-source citations and materially better hypothesis quality at the synthesis step. Single-tool workflows are fine for exploratory questions; multi-tool chains are for anything a partner will sign off on.
Is McKinsey Lilli available to non-McKinsey consultants?
No. Lilli is an internal-only platform, as are Bain Sage and BCG Deckster. External consultants have to build the equivalent using commercial tools — which is exactly the workflow this post describes. The good news is that the commercial stack in 2026 is essentially at parity with the internal firm tools for external research; the gap is on internal corpus access, which is what Glean solves for large corporate strategy teams.
How do I handle client confidentiality when using these tools?
Three rules. First, only use enterprise-tier subscriptions with signed data processing agreements. Second, verify your firm's IT and compliance policy explicitly permits the tool — many banks, healthcare clients, and government engagements prohibit third-party AI processing without a specific carve-out. Third, when in doubt, redact identifying information before uploading. Never paste raw client data into a consumer-tier account.
What's the difference between Perplexity Deep Research and Perplexity Enterprise?
Perplexity Deep Research is a feature — a multi-step agent mode available inside any Perplexity plan (including Pro at $20 per month). Perplexity Enterprise is the tier that adds admin controls, workspace-level SSO, data protection commitments, higher rate limits, and access to enterprise-only models. For consulting work with any client-sensitive context, Enterprise is the correct choice.
How long does it take to become fluent with this deep research agent workflow?
For a consultant who already uses ChatGPT or Claude daily, becoming productive with the chain takes roughly 10 to 15 hours of deliberate practice across three or four real engagements. Fluency — the point where you can adjust the chain on the fly to match a novel question — takes more like 40 to 60 hours over 3 to 4 months. Most of the learning is in prompt design and knowing which tool to skip.
Do partners at MBB firms actually use these tools themselves, or is it staff-only?
Both, but the split is heavily staff-weighted. Partners typically consume the outputs — memos, decks, audio overviews — rather than orchestrate the chain themselves. That said, the fastest-adopting partners at Bain and McKinsey are using Claude Projects and NotebookLM directly for their own thinking and prep. Expect this to shift toward partner-level fluency over the next 18 months.
Can this workflow replace primary research like expert interviews?
No, and anyone selling it as a replacement is oversimplifying. The workflow dramatically compresses the desk-research phase, which frees up budget and time for higher-quality primary research — more expert calls, deeper site visits, better survey design. The best case teams in 2026 use AI to buy back time for the primary work that actually creates defensibility.
What's the biggest hidden cost of adopting this deep research agent workflow?
Prompt library maintenance. The tools themselves are cheap. What's expensive is the time to build, tune, and version a library of prompts that consistently produce partner-quality output. Teams that treat prompts as throwaway inputs plateau at mediocre results. Teams that treat prompts as versioned assets, reviewed and refined every quarter, compound their advantage over 12 to 18 months.
How do I keep this stack from going stale as models change?
Two habits. First, rerun a benchmark task — the same due-diligence prompt across all four tools — once a quarter and update your default-tool routing based on the results. Second, subscribe to two or three high-signal AI newsletters (Interconnects, Latent Space, and Anthropic's own release notes) rather than trying to track every model release. The tool routing matters more than the model version.
The Bottom Line
The deep research agent workflow is not a single tool. It is a chain — Perplexity for external scans, Glean for internal precedent, ChatGPT for financial context, Claude for synthesis, and NotebookLM for briefing artifacts — with human judgment at the handoffs. The firms that will not publish it are hiring aggressively for people who can run it. The consultants who master it in 2026 will bill the same hourly rate on 60% of the desk-research time, or bill the same hours on materially harder problems. Either way, the delta compounds. Start with one prompt from the anatomy section above, on your next real engagement. Not a practice run. Real work, real stakes. Then the second prompt on the engagement after that. Fluency arrives faster than most people expect.






