The Continuity Brief

Tabletop DR Exercises for Small IT Teams

Lean IT teams need disaster drills designed for their actual constraints.

Staff Writer · · 11 min read
Cover illustration for “Tabletop DR Exercises for Small IT Teams”
Disaster Recovery Planning · September 17, 2026 · 11 min read · 2,436 words

Small IT teams fail during disasters because the plan assumes a staffing model that doesn't exist. They fail because the plan assumes a staffing model that doesn't exist, with a crisis lead here, a communications officer there, and a legal team on standby. A tabletop exercise built around that fantasy roster wastes the one afternoon a lean team actually has to prepare. The fix is a tabletop designed from the ground up for the constraints a three-person IT shop actually lives with. It's a tabletop designed from the ground up for the constraints a three-person IT shop actually lives with.

The stakes aren't abstract. According to Mitratech, more than half of IT and data center outages cost over $100,000, and 16% run past $1 million. Those numbers don't discriminate by company size. A five-person IT team at a mid-sized firm faces the same downtime math as a large enterprise data center, just without the redundant staff to absorb the hit.

What a DR tabletop exercise is, and what it is not

A tabletop exercise is a discussion. Team members sit down, walk through a scenario someone has designed in advance, and talk through how they'd respond, step by step. No systems get touched. Nothing fails on purpose. The entire exercise lives in conversation, notes, and the occasional awkward silence when someone realizes they don't actually know who has the credentials to the backup server.

That's different from a live simulation, which runs against real systems and data flows under production-like conditions, or a failover test, which actually shifts workloads to a backup environment and clocks recovery times against RTO and RPO targets. Those exercises measure infrastructure. A tabletop measures something else entirely: decision-making, role clarity, the order in which people communicate, and whether the sequence of operations in someone's head matches what's written down.

For a small team, that distinction is the whole point. Running a tabletop doesn't require a second data center, a sandboxed test environment, or a scheduled outage window. It requires a room (or a call), a scenario, and one to three hours, depending on how complicated the scenario gets. That's a realistic ask even for a team of two or three people who don't have a spare afternoon, let alone a spare data center.

None of this means the exercise is a substitute for the real thing. A tabletop won't produce the infrastructure measurements a failover test will, and it can't replicate the adrenaline and noise of an actual 2 a.m. outage call. Some physical or infrastructure constraints become visible only once systems are actually live, no matter how carefully the scenario gets written, though a well-designed discussion can catch a surprising number of them before that happens. Still, for a team with no exercise history at all, a tabletop is the lowest-friction way in. It delivers real preparedness value without demanding resources the team doesn't have.

What a well-run tabletop produces for a small team

The output is a shared, tested understanding of who does what, in what order, when a specific named system goes down. It's a shared, tested understanding of who does what, in what order, when a specific named system goes down, not a theoretical org chart, but a decision path the team has actually walked through out loud.

Consider what surfaced during a tabletop exercise run by University of Michigan LSA Technology Services. A researcher's grant required a particular server, and the data on it, to stay physically in a locked rack; the team wasn't permitted to move it to another data center or even another rack during a disaster recovery event. Not everyone on the team knew that constraint existed until the tabletop exercise put it on the table. Participants later described the roundtable format as useful for "brain-storming" and said it helped people not directly involved in running the service understand what a disaster would actually mean for it.

That's the value proposition in miniature: gaps that live in one person's head get pulled into the open before an actual outage forces the issue. Beyond that specific example, a well-run session tends to expose single points of failure across people, process, and technology, each with a name attached to the fix. It validates or invalidates assumptions buried in runbooks, contact trees, and vendor SLAs that nobody has actually reread in a year. It clarifies who decides, who acts, who informs, and, critically, who covers when the usual person is out sick or on a plane. And it leaves behind documentation, notes, timelines, outcomes, that double as evidence controls actually exist, which matters the next time an auditor asks.

There's also an education function that's easy to undervalue. Mark Lance, VP of DFIR and Threat Intel at GuidePoint Security, has observed that senior leadership teams walking through a ransomware scenario for the first time "walk away saying, 'I know more about this and the potential risks.'" The same dynamic plays out lower down the org chart: a junior sysadmin who's never thought about what recovering a failed storage appliance actually involves comes out of the exercise with a mental model they didn't have going in.

Scale differs, but the pattern holds even at the top of the field. ENISA's pan-EU Cyber Europe 2024 exercise, run in June 2024 across the energy sector with 30 countries and over a thousand participants, found that more than 90% of participants felt more prepared to handle cyber incidents afterward. A three-person IT shop running its own tabletop isn't operating at that scale, but the underlying mechanism, rehearsal building confidence and competence, is the same one at work.

If the team wants to know whether the program is actually working, it should track how many exercises get run, how many gaps each one surfaces, how many corrective actions get assigned, how long those actions take to close, and whether last quarter's fixes were actually in place before this quarter's exercise started.

The six steps of a small-team tabletop, adapted for lean operations

The standard structure, set goals, pick participants, design the scenario, run it, review it, define improvements, holds up fine. What changes for a small team is how each step gets executed.

Establishing goals and scope means tying the exercise to one or two critical business functions, not the entire infrastructure at once. A small team can't usefully discuss everything in a single session, and trying to will produce a shallow conversation about everything instead of a deep one about anything. Decide what's explicitly out of scope before starting, and set time limits per phase so the discussion doesn't sprawl.

Selecting participants means keeping the group small, but making sure it covers the essential functions: someone with the authority to make priority calls under pressure, someone who actually knows the affected systems, and a scribe to capture decisions and timestamps as they happen. Name backups for key roles now, before the exercise starts, so "what happens if this person is unavailable" gets built into the scenario rather than quietly avoided. If a vendor's platform or SLA sits at the center of the scenario, consider inviting them. In a genuinely tiny team, the person facilitating may also need to participate; that's a real limitation, and it's better addressed with a written scenario card than pretended away.

Designing the scenario and materials means assembling the artifacts before the session starts: current contact lists, network and application diagrams, existing runbooks, vendor SLA documents. Hunting for these mid-exercise burns time that should go toward the actual discussion. Check the scenario's realism with whoever knows the systems best before running it in front of the full group, and build in complexity gradually through injects, new information dropped in partway through, rather than dumping everything on the table at once.

Running the exercise means the goal is to learn something, not to look competent. Design the scenario to be genuinely hard: missing information, conflicting priorities, an inconvenient surprise halfway through. Track injects, decisions, owners, and timestamps on a shared board as the exercise unfolds. The facilitator's job is to keep the pace moving and press on decision points, not to solve them for the team.

Reviewing, or running the hot wash, should happen immediately afterward, while the details are still sharp. Document what worked, what broke down, and where the team got stuck, not just a list of what went wrong, but where the group hesitated or disagreed.

Defining improvements and assigning owners means every gap identified needs a name, a specific action, and a target date attached before the session ends. Writing up a short after-action report means covering an overview, the scenario description, key strengths and weaknesses, specific recommendations, and an implementation timeline. Then schedule the next exercise before closing this one out. Keeping the program going matters more than making any single session perfect.

How often? Regularly enough that the program stays current, and any time something material changes, a new hire, a new system in production, or a past incident that revealed a gap nobody had planned for.

Five scenarios worth running, and what each one tests for small teams

Pick scenarios based on the team's actual crown jewels, the systems and data whose failure would do the most damage, not a generic list of threats copied from an enterprise template.

A primary server crash simulates complete failure of on-premise servers, forcing failover to a secondary site or cloud backup. It tests whether everyone actually knows where the backups live, who holds the credentials, and in what order systems need to come back online, and whether the runbook still matches reality. The pressure point for small teams: the person who built the backup system in the first place is often the one who's unreachable when it matters.

A ransomware attack involves encrypted systems, a ransom demand, possibly data exfiltration, and a decision about whether to pay or restore from backup. This tests isolation procedures, backup integrity, how leadership and affected customers get notified, and legal obligations around disclosure. Communication and notification paths are often completely undefined until a tabletop forces the question into the open. CISA's Tabletop Exercise Packages include ransomware scenarios and cost nothing to use.

A cloud provider outage occurs when a major SaaS platform or cloud provider goes down, and everything downstream of it stops working. This tests whether the team has documented manual fallback procedures and knows the escalation path with the vendor's SLA. Small teams often lean hard on a handful of cloud services, which means one outage can disable a disproportionate share of daily operations, and many haven't mapped out exactly how dependent they are until this scenario makes them do it.

A third-party or supply chain breach enters through a vendor's systems, and the team has to figure out scope and containment without any control over what the vendor does on their end. This tests vendor contact procedures, the ability to isolate third-party integrations quickly, and whether anyone actually knows what data that vendor had access to. Small IT teams tend to have thinner vendor management processes than larger shops, which makes this gap hard to spot without running the exercise. Sygnia's investigation found an AI-assisted cloud intrusion that moved from initial access to broad compromise in just 72 hours, a pace that makes vendor response time a live variable, not a footnote.

Key-person unavailability during an active incident. The system owner or IT lead is unreachable the moment an outage starts, and the rest of the team has to run recovery without them. This tests documentation completeness, whether credentials are actually accessible to more than one person, and whether anyone besides the usual expert can perform the recovery steps. This is the scenario most specific to lean teams, and the one most enterprise playbooks skip entirely. University of Michigan LSA Technology Services ran exercises on individual services, including authentication, storage appliances, build automation, and Windows web hosting, precisely to surface this kind of per-system dependency before it became a crisis.

Teams with AI tools in production should think about adding a related scenario or inject: an LLM leaking data it shouldn't, an AI agent reaching into a system beyond its intended scope, or an attacker using AI to move faster than a human-paced response plan expects. Enrique Alvarez at Google Cloud has pointed to threat actors exploiting CVEs at an increased rate with AI assistance, and Sygnia's 72-hour compromise timeline makes the point concrete. Small teams don't need to build a whole scenario around this. A single inject, "the attacker moved faster than expected, here's the updated timeline", can reveal whether the team's response pace still assumes a pre-AI threat model.

CISA offers over 100 free Tabletop Exercise Packages covering ransomware, insider threats, phishing, and ICS compromise. For a team with no budget for this kind of exercise design, it's a practical place to start rather than build from scratch.

Common failure modes in small-team tabletops and how to avoid them

The exercise turns into a status meeting. Participants start describing what they'd normally do in vague, general terms instead of actually working through the specific decisions the scenario demands. The fix is for the facilitator to keep dropping injects, new information that shifts the situation mid-conversation, so the group is forced into real decisions instead of reciting policy they already know by heart.

Only the loudest or most senior person talks. The engineer who actually maintains the system in question goes quiet while the manager answers every question on their behalf. Direct scenario questions at specific roles, not at the room in general. The person who runs the backup should walk through the restoration steps themselves, not have someone else narrate it for them.

If the scenario is too easy and the team sails through every decision point without friction, the exercise has confirmed what everyone already knew instead of finding what nobody had thought about. Build in missing information, conflicting priorities, and resources that aren't available when needed. The value of the exercise lives in the questions the team can't answer cleanly, not the ones they can.

Physical and contractual constraints stay invisible: rack locations, grant restrictions, and vendor contract terms don't appear in a generic scenario unless someone specifically asks about them. The University of Michigan example is instructive here precisely because that constraint wasn't hypothetical: it was a real condition attached to a real piece of equipment, and a structured discussion brought it to light. Scenario design should include a deliberate prompt, something like "are there any physical, contractual, or legal constraints on how this system can be recovered", rather than assuming those questions will come up on their own.

Sources

  1. What is a Disaster Recovery Tabletop Exercise?
  2. How to Run Incident Response Tabletop Exercises in 2025 | Sygnia
  3. CISA Tabletop Exercise Packages | CISA
  4. Disaster Recovery Plan (DRP) Tabletop Exercises | U-M LSA LSA Technology Services

More in Disaster Recovery Planning