Back to Blog
Software Testing

Mock vs. Stub vs. Fake in 2026: Understanding Test Doubles and When to Use Each in Modern Test Suites

Avanish Pandey

October 11, 2026

Mock vs. Stub vs. Fake in 2026: Understanding Test Doubles and When to Use Each in Modern Test Suites

Mock vs. Stub vs. Fake in 2026: Understanding Test Doubles and When to Use Each in Modern Test Suites

Mock vs Stub vs Fake test doubles guide

Mock, stub, and fake are three distinct types of test doubles that serve different purposes in unit and integration test suites. A stub returns predetermined data so a test can exercise a code path without calling a real dependency. A mock records how it was called and asserts that specific interactions occurred — it tests behavior, not just output. A fake is a working implementation with simplified logic, suitable for use across an entire test suite where stubs would be too granular and a real dependency too expensive. Choosing the wrong type produces tests that pass but do not catch the bugs they were designed to catch, or tests that couple too tightly to implementation details and fail on every refactor. Astaqc's test automation services include test architecture review to ensure teams use the correct test double type for each dependency and test scope. The complete software testing guide covers test double strategy as part of foundational unit testing practice.

What Is a Test Double and Why the Taxonomy Matters

The term test double was introduced by Gerard Meszaros in xUnit Test Patterns to describe any object that stands in for a real dependency in a test. The taxonomy — dummy, stub, spy, mock, fake — gives names to meaningfully different patterns that serve different verification goals. In practice, many developers use "mock" as a generic term for any test double, which obscures the distinction between testing state (did the code produce the right output?) and testing behavior (did the code call its dependencies in the right way?). Blurring this distinction leads to tests that verify the wrong thing.

The 2026 landscape makes the taxonomy more relevant, not less. AI-assisted code generation tools frequently produce tests using mocking frameworks without reasoning about which test double pattern is appropriate for the dependency being replaced. A generated test that mocks every collaborator produces a test that validates only the internal wiring of the unit under test — it cannot catch bugs that emerge from the interaction between real components. Teams that understand the taxonomy can review generated tests critically and identify which doubles add verification value and which add only coupling.

The practical distinction that matters most in daily engineering work is between mocks and stubs. A stub is passive: it returns a value when asked. A mock is active: it records interactions and verifies that the code under test called it in the expected way. Confusing the two means either missing a verification (using a stub where a mock would catch a missing method call) or over-specifying (using a mock where a stub would allow the implementation to change without the test failing). Astaqc's software testing services include unit test audits that identify over-mocked tests as a form of test debt. The manual vs. automated testing guide covers when unit tests with doubles provide enough coverage versus when integration or end-to-end tests are required.

Stubs: Controlling Inputs to the Code Under Test

A stub replaces a dependency and returns a controlled value when the code under test calls it. The test's assertion is on the output or state of the code under test — not on how the stub was called. A payment processing function that calls a currency conversion API should be tested with a stub that returns a fixed exchange rate, so the test can assert that the conversion arithmetic produces the correct result without making a network call.

Stubs are the most common test double and the most appropriate choice when the dependency provides data that the code under test processes. In 2026, stubs are used extensively in AI-augmented pipelines where LLM API calls are replaced with stubs that return deterministic responses, allowing the surrounding application logic to be tested without incurring API costs or dealing with non-deterministic model output. The stub's job is to eliminate variability so the test can focus on the logic under test.

The characteristic error with stubs is returning values that do not represent realistic scenarios. A stub that always returns success does not test the error handling path. A stub that returns data in a format different from the real API response produces tests that pass in isolation but fail in integration. The fix is to define stub fixtures from real API responses (captured and stored as JSON fixtures) rather than hand-crafting stub return values. This way the stub contracts are anchored to actual dependency behavior and update when the API changes. Astaqc's manual testing services include exploratory testing that catches the edge cases that stub-based unit tests miss by their design. The AI in software testing guide covers how stubs are used specifically in LLM application testing.

Mocks: Verifying Behavior and Interaction

A mock is a test double that records interactions and verifies them as part of the test assertion. Where a stub makes assertions about output (what the code returned), a mock makes assertions about behavior (how the code called its dependencies). A notification service that is supposed to send an email when an order is placed should be tested with a mock — the test asserts that the email service's send method was called once, with the correct recipient and subject. If the code under test never calls the email service, the mock assertion fails even if the order record was created correctly in the database.

Mocks are the appropriate choice when the interaction with the dependency is itself part of the contract being tested. Side-effect dependencies — sending email, writing to a queue, logging to a monitoring service — have no output to assert on from the caller's perspective. A stub that returns success for these calls would produce a passing test even if the code never called the dependency at all. A mock catches the missing call.

The characteristic error with mocks is over-specifying. A test that mocks a repository and asserts that findByIdWithEagerLoadedRelations was called with a specific SQL-level parameter is testing the implementation, not the contract. When the code is refactored to use a different query method, the mock assertion fails even though the behavior is unchanged. The fix is to mock at the boundary the test actually cares about — usually the interface or public API of the collaborator, not the internal implementation details. In modern TypeScript and Python codebases with interface-driven design, this means mocking the interface rather than the concrete class. Astaqc's hire QA team services include test suite reviews that identify over-mocked tests and refactor them toward interface-level assertions. The regression testing guide covers how fragile mocks are a leading cause of false failures in regression suites.

Fakes: Working Implementations for Entire Test Suites

A fake is a simplified but working implementation of a dependency. Unlike stubs and mocks, which are configured per-test, a fake can be shared across an entire test suite because it maintains state and behaves correctly across multiple calls. An in-memory database is the canonical fake — it implements the same interface as the real database (create, read, update, delete), maintains state between calls within a test, and resets between tests. A test suite that uses an in-memory database fake can run without a real database server and without per-test stub configuration.

Fakes become valuable when stubs become too granular to maintain. A service with 20 methods, each of which needs to be stubbed differently in every test, creates a maintenance burden that grows with the number of tests. Replacing the stub with a fake that implements the full interface once means each test interacts with a working implementation rather than a per-test stub configuration. The fake pays a one-time implementation cost and reduces per-test setup code at scale.

TypeWhat it doesWhat the test asserts onWhen to use
StubReturns a fixed valueOutput of the code under testDependency provides input data
MockRecords interactions and verifies themHow the code called the dependencyDependency has side effects
FakeWorking simplified implementationOutput or state, like real dependencyFull test suite, stateful dependency
SpyWraps real implementation, records callsInteractions on a real or fake dependencyVerify calls on a real implementation
DummyPlaceholder that is never actually usedNothing — filler for required parametersRequired parameter with no test relevance

The risk with fakes is drift from the real implementation. An in-memory database fake that does not enforce uniqueness constraints will pass tests that fail against a real database with a unique index. Keeping fakes aligned with real dependency behavior requires the same discipline as maintaining any shared test infrastructure. Contract tests — tests that run against both the fake and the real dependency to verify they behave identically on the scenarios the fake covers — are the standard mitigation. Astaqc's performance testing services cover scenarios where fakes cannot substitute for real dependencies — load and throughput characteristics that only emerge at production scale require real infrastructure.

Frequently Asked Questions

Is it wrong to use the word "mock" to mean any test double?

It is imprecise but common. In casual conversation, "mock" often means any test double, and most engineers understand it in context. The precision matters when reviewing or writing tests, because choosing the wrong double type produces tests that either miss real bugs or break on refactors. In code reviews and test design discussions, using the precise terms — stub when returning data, mock when verifying interactions, fake when providing a working implementation — leads to clearer reasoning about what a test is actually verifying.

When does a spy make sense over a mock?

A spy wraps a real or working implementation and records method calls, allowing assertions on interactions without replacing the real behavior. A spy is appropriate when the test needs to verify that a method was called (like a mock) but also needs the real implementation to run (unlike a mock, which replaces the real behavior). Logging or metrics recording are typical spy candidates — the test verifies that the log was called with the right message while also letting the real computation complete so its output can be asserted.

Can a test use multiple types of test doubles simultaneously?

Yes, and most tests do. A unit test for an order processing function might use a stub for the pricing service (provides input data), a mock for the email service (verifies the confirmation email was triggered), and a dummy for a logger parameter (required by the constructor but irrelevant to the test). Each double plays a different role in the same test. The choice per dependency follows the same rule: is the dependency providing data to the code under test (stub), receiving a side-effect from the code under test (mock), or irrelevant to the test (dummy)?

How do test doubles relate to dependency injection?

Dependency injection is the mechanism that makes test doubles possible. If a class creates its own dependencies internally (using new in the constructor or calling a static method), there is no seam to insert a test double. Injecting dependencies through the constructor or a factory parameter allows tests to pass in stubs, mocks, or fakes instead of real implementations. This is why dependency injection and testability are treated as related concerns in clean architecture guidance — design for injection is design for testability. Refactoring legacy code toward dependency injection is one of the most reliable paths to increasing unit test coverage on previously untestable code.

How does this taxonomy apply to testing AI-integrated applications?

AI-integrated applications have dependencies with characteristics that do not fit neatly into traditional test double patterns. LLM API calls are non-deterministic, slow, and expensive — stubs that return fixed responses are the correct approach for unit testing the application logic that processes LLM output. The challenge is that LLM output format and content variation is much wider than a traditional API response, so stubs need to cover the relevant output variations (structured vs. unstructured, success vs. refusal, correct format vs. hallucinated schema). Contract testing between stubs and real API calls on a representative sample remains the standard for catching stub drift. See Astaqc's AI in software testing guide for coverage strategies specific to AI-integrated applications, and the TestInspector product page for a no-code test automation tool that handles dynamic, AI-augmented web application testing without requiring manual test double configuration.

The test double you choose determines whether your test validates behavior or just confirms that code ran. Mocks, stubs, and fakes are not interchangeable — each answers a different question about your system.

Avanish Pandey

October 11, 2026

icon
icon
icon

Subscribe to our Newsletter

Sign up to receive and connect to our newsletter

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Latest Article

Ask our AI assistant…