1 / 20

Fewer Tests, More Confidence

A Rails Testing Philosophy for the AI Era

Ifat Ribon

Rocky Mountain Ruby | 09.2026

AI can generate tests fast...

...but it's not free

More to review
More to reason about
More to maintain
More to wait for in CI
More money
More, more, more

AI can generate more testsbut not always better tests

Define a Testing Philosophy

TESTING PHILOSOPHY

Why We Test

1
Validate code meets requirements
2
Protect essential user flows & business logic
3
Catch key failure states & critical interdependencies
4
Enable refactoring with confidence

Tests document requirements...and now they serve as context

TESTING PHILOSOPHY

Test Categories

Unit
Logic in isolation
Best for model specs and edge cases
FAST & FREQUENTLY
System
User experience, integrated components
Best for happy paths in key user flows
MODERATE & SELECTIVELY
Integration
How components fit together
Best for external dependencies and complex flows
SLOW & PERIODICALLY

TESTING PHILOSOPHY

The Test Pyramid

UNIT
Fast · Cheap
SYSTEM
Moderate
INTEGRATION
Slow · Expensive
Less Granular
More Expensive
More Granular
Less Expensive

TESTING PHILOSOPHY

Testing Principles

01
Test behavior,
not implementation
02
Keep tests
deterministic
03
Balance coverage
with cost
04
Use good test design to
promote good code design

TESTING PHILOSOPHY

Testing Principles

01
Test behavior,
not implementation
02
Keep tests
deterministic
03
Balance coverage
with cost
04
Use good test design to
promote good code design

TESTING PHILOSOPHY · PRINCIPLE 01

Test Behavior, Not Implementation

Avoid testing library / language code
Focus on public methods and their results
Assert tests pass after refactors
❌ Implementation-coupled
test "calculates the discount" do order = Order.new(total: 100) assert_equal 10, order.send(:calculate_discount) end
✅ Behavior-focused
test "applies discount to eligible order" do order = Order.new(total: 100, promo: "SAVE10") assert_equal 90, order.discounted_total end
A good test survives refactors

TESTING PHILOSOPHY

Testing Principles

01
Test behavior,
not implementation
02
Keep tests
deterministic
03
Balance coverage
with cost
04
Use good test design to
promote good code design

TESTING PHILOSOPHY · PRINCIPLE 02

Keep Tests Deterministic

Keep data local
Stub collaborators in unit/system tests
Control time when it affects behavior
❌ Flaky — time-dependent
test "trial is active during the window" do user = User.create!(trial_ends_at: 3.days.from_now) assert user.trial_active? end
✅ Deterministic — frozen time
test "trial is active during the window" do travel_to Time.zone.parse("2026-06-01 12:00") do user = User.create!(trial_ends_at: "2026-06-04 12:00") assert user.trial_active? end end
A failing test means something is broken — not "try re-running"

TESTING PHILOSOPHY

Testing Principles

01
Test behavior,
not implementation
02
Keep tests
deterministic
03
Balance coverage
with cost
04
Use good test design to
promote good code design

TESTING PHILOSOPHY · PRINCIPLE 03

Balance Coverage with Cost

Choose test types intentionally
Minimize data setup
Avoid testing external dependencies directly
❌ System test for pure logic
test "applies promo discount" do visit new_order_path fill_in "Promo code", with: "SAVE10" click_on "Apply" assert_text "$90.00" end
✅ Unit test — same confidence, far cheaper
test "applies promo discount" do order = Order.new(total: 100, promo: "SAVE10") assert_equal 90, order.discounted_total end
Reach for the cheapest test that gives confidence

TESTING PHILOSOPHY

Testing Principles

01
Test behavior,
not implementation
02
Keep tests
deterministic
03
Balance coverage
with cost
04
Use good test design to
promote good code design

TESTING PHILOSOPHY · PRINCIPLE 04

Use Good Test Design

Tests should have a single reason to fail
Results should be objective and discrete
Failures are actionable and specific
❌ Multiple reasons to fail
test "user signup" do user = User.create!(name: "Jo", email: "jo@ex.co") user.sign_up! assert user.confirmed? assert_equal "free", user.plan assert user.welcome_email_sent? end
✅ Single reason to fail
test "new user gets the free plan" do user = User.create!(name: "Jo", email: "jo@ex.co") user.sign_up! assert_equal "free", user.plan, "Expected new signups to default to free plan" end
A convoluted test points to convoluted code

Rein in the Agent

HARNESS THE AGENT

AI Test Flow

Generate
AI
→
Review
Human
→
Prune
Collab
→
Harden
Collab
→
Capture
AI
order_test.rb — 14 tests generated 🤖
test "applies promo discount to eligible order" do order = Order.new(total: 100, promo: "SAVE10") assert_equal 90, order.discounted_total end test "discount never drops total below zero" do order = Order.new(total: 5, promo: "SAVE100") assert_equal 0, order.discounted_total end test "pending? returns true when status is pending" do order = Order.new(status: :pending) assert order.pending? end test "discounted_total returns a reduced total" do order = Order.new(total: 100, promo: "SAVE10") assert_operator order.discounted_total, :<, 100 end

HARNESS THE AGENT

AI Test Flow

Generate
AI
→
Review
Human
→
Prune
Collab
→
Harden
Collab
→
Capture
AI

HARNESS THE AGENT

Reviewing Tests

00Does it do what it says?
The description is the spec. The body is the proof.
01Behavior
Tests library, framework, or language code instead of custom logic
Tests private methods or implementation details
Would break on an internal refactor that doesn't change behavior
02Determinism
Relies on data outside the system under test
Hits real collaborators (HTTP, queues, mailers) in a unit or system test
Depends on execution order or shared mutable state
03Balance
Uses a heavier test type than the behavior requires
Creates superfluous data that isn't needed for the assertion
Hits external dependencies directly in a unit or system test
04Design
Has multiple reasons to fail
Produces ambiguous or non-discrete results
Fails with a message that doesn't tell you what broke or why
order_test.rb — keep or delete? 🔍
# ✅ Core behavior — keep test "applies promo discount to eligible order" do order = Order.new(total: 100, promo: "SAVE10") assert_equal 90, order.discounted_total end # ✅ Edge case worth preventing — keep test "discount never drops total below zero" do order = Order.new(total: 5, promo: "SAVE100") assert_equal 0, order.discounted_total end # ❌ Tests Rails' enum machinery, not your app — delete test "pending? returns true when status is pending" do order = Order.new(status: :pending) assert order.pending? end # ❌ Redundant with the first test — delete test "discounted_total returns a reduced total" do order = Order.new(total: 100, promo: "SAVE10") assert_operator order.discounted_total, :<, 100 end

HARNESS THE AGENT

AI Test Flow

Generate
AI
→
Review
Human
→
Prune
Collab
→
Harden
Collab
→
Capture
AI

HARNESS THE AGENT

Make the Suite Lean

Cut the noise!
Delete what failed review: tests of the framework, redundant duplicates, brittle one-offs.
Merge the dupes!
Consolidate overlapping tests. Extract shared setup into helpers and factories.
Shape what stays!
Refactor the keepers to match suite conventions. Tests are code, so hold them to the same bar.
Be ruthless about deleting low-value, brittle, or redundant AI-generated tests!
order_test.rb — pruning the suite ✂️
# ✅ Core behavior — keep test "applies promo discount to eligible order" do order = Order.new(total: 100, promo: "SAVE10") assert_equal 90, order.discounted_total end # ✅ Edge case worth preventing — keep test "discount never drops total below zero" do order = Order.new(total: 5, promo: "SAVE100") assert_equal 0, order.discounted_total end # ❌ Tests Rails' enum machinery, not your app — delete test "pending? returns true when status is pending" do order = Order.new(status: :pending) assert order.pending? end # ❌ Redundant with the first test — delete test "discounted_total returns a reduced total" do order = Order.new(total: 100, promo: "SAVE10") assert_operator order.discounted_total, :<, 100 end

HARNESS THE AGENT

AI Test Flow

Generate
AI
→
Review
Human
→
Prune
Collab
→
Harden
Collab
→
Capture
AI

HARNESS THE AGENT

Fill the Gaps

Hunt for missing cases!
Find edge cases the suite doesn't cover, e.g., nil, empty, boundary values, unauthorized access.
Generate variants!
Same behavior, hostile inputs, e.g., zero totals, expired promos, duplicate submissions, missing records.
Prove it catches bugs!
Introduce plausible bugs into the code. Every one the suite misses is a test you're missing.
The pruned suite is lean, now make sure it's complete!
order_test.rb — filling the gaps 🛡️
# ✅ Core behavior — keep test "applies promo discount to eligible order" do order = Order.new(total: 100, promo: "SAVE10") assert_equal 90, order.discounted_total end # ✅ Edge case worth preventing — keep test "discount never drops total below zero" do order = Order.new(total: 5, promo: "SAVE100") assert_equal 0, order.discounted_total end # 🛡️ Edge case added — unknown code test "unknown promo applies no discount" do order = Order.new(total: 100, promo: "BOGUS") assert_equal 100, order.discounted_total end # 🛡️ Edge case added — missing promo test "nil promo leaves total unchanged" do order = Order.new(total: 100, promo: nil) assert_equal 100, order.discounted_total end

HARNESS THE AGENT

AI Test Flow

Generate
AI
→
Review
Human
→
Prune
Collab
→
Harden
Collab
→
Capture
AI

HARNESS THE AGENT

Bolster the AI Harness

Rules
Codify hard constraints and convetions: what to avoid, what patterns are banned, how to review its output.
RULE.md
# Testing Rules ## Principles 1. **Test behavior, not implementation.** Test public methods and their observable results. Do not test Rails internals, Ruby stdlib, or gem code. Tests must not break when internals are refactored. 2. **Keep tests deterministic.** Data is local to each test. Stub external collaborators. Control time with `travel_to`. Never depend on record ordering without an explicit `ORDER BY`. A failing test means something is broken — not "try re-running." 3. **Balance coverage with cost.** Choose the cheapest test type that gives confidence: unit for logic, system for user flows, integration for external contracts. Minimize data setup — only create data the assertion requires. Stub HTTP/APIs/filesystem in unit and system tests. 4. **Good test design promotes good code design.** One reason to fail per test. Objective pass/fail. Actionable failure messages. If a test is hard to write, simplify the code under test. ## Review checklist Start with the gate. Then check the twelve red flags. Every test — written or generated — must pass the gate and be clean of all twelve flags. If any flag is true, fix or delete the test. ### 00 · Does it do what it says? The test name makes one claim about behavior. The assertions prove that claim and nothing else. ### 01 · Behavior - Tests library, framework, or language code instead of custom logic - Tests private methods or implementation details - Would break on an internal refactor that doesn't change behavior ### 02 · Determinism - Relies on data outside the system under test - Hits real collaborators (HTTP, queues, mailers) in a unit or system test - Depends on execution order or shared mutable state ### 03 · Balance - Uses a heavier test type than the behavior requires - Creates superfluous data that isn't needed for the assertion - Hits external dependencies directly in a unit or system test ### 04 · Design - Has multiple reasons to fail - Produces ambiguous or non-discrete results - Fails with a message that doesn't tell you what broke or why ## Conventions - `Minitest::Test` for units, `ActionDispatch::SystemTestCase` for system tests. - Name methods `test_<description_of_behavior>` — should read as a sentence. If the name has "and" in it, split the test. - Prefer `assert_equal`, `assert_nil`, `assert_includes`, `assert_raises`, `assert_difference` over bare `assert`.

HARNESS THE AGENT

Bolster the AI Harness

References
Give AI documentation and examples: the docs, guides, and example tests the agent should draw from.
references/unit_tests.md
## Unit Test Best Practices Unit tests go in `test/models/`, `test/services/`, `test/jobs/`, `test/mailers/`, or the appropriate directory mirroring `app/`. ```ruby require "test_helper" class WidgetTest < ActiveSupport::TestCase setup do # minimal shared context — only what multiple tests need end test "calculates total price from quantity and unit price" do widget = Widget.new(quantity: 3, unit_price: 10_00) assert_equal 30_00, widget.total_price end end ``` When using fixtures or factories, build only the attributes relevant to this test. When the method under test calls an external service, API client, or mailer, stub that call. When behavior depends on time, wrap the test in `travel_to`.

HARNESS THE AGENT

Bolster the AI Harness

Skills
Teach AI how to write and maintain the suite: point it to the tools and conventions it should use.
SKILL.md
--- name: test description: Generate and maintain Minitest tests for a Rails application. Use this skill whenever the user asks to write tests, add test coverage, review existing tests, fix a failing test, or mentions testing in the context of a Rails + Minitest codebase. --- # Testing This app uses minitest and relies on fixtures over factories. In this app the following semantics are used: 1. Unit tests - isolated and discrete business logic. 2. System tests - user behavior and experience, capybara tests 3. Integration tests - Third-party integrations and complex interdependencies (full backend workflows, including database triggers) There are several references on best practices for writing minitests, and strong test coverage in general. These are captured in the references folder and listed below. These *must* be read before writing, reviewing, or refactoring any test: 1. [Testing Rules](../../rules/test-rules.md) 2. [Minitest Rules](../../rules/minitest-rules.md) 3. [Test Style Guide](../../rules/test-style-guide.md) Tests for specific areas of the app should leverage the below references. Review each one if you are attempting tests in a given area: 1. [Unit Tests](./references/unit_tests.md) 1. [Model Tests](./references/model_tests.md) 1. [Controller Tests](./references/controller_tests.md) 1. [System Tests](./references/system_tests.md) ## Before writing any test 1. **Read the references** if you haven't this session. The rules there are authoritative. 2. **Identify the behavior under test.** State it in one sentence. If you can't, the scope is too broad, split it. 3. **Choose the right test type.** Logic in isolation → unit test. User-facing flow → system test. External contract → integration test. 4. **Check for redundancy.** Scan existing tests for the file or class. Do not duplicate coverage that already exists. ## Running the Tests - Unit tests: `bin/rails test` - System tests: `bin/rails test:system` - Integration tests: `bin/rails test:integration`. - Single file with `bin/rails test test/models/user_test.rb`, or single example with `...:42`, ## How to maintain the suite Tests accumulate debt like any other code. When the codebase evolves: - **Extract shared setup** into helpers (`test/support/`) or factory methods when the same creation logic appears in three or more files. - **Replace repeated inline stubs** with a shared helper method. - **Rename tests** whose names no longer match the behavior they verify. - **Refactor tests alongside the code they test.** When a model changes shape, update its tests in the same PR. - **Delete aggressively.** Tests that cover dead behavior, test framework code, or duplicate another test are liabilities. Remove them and note why in the PR. ## How to explain your work When you generate tests, briefly note in your summary what each test covers and why it's worth having. This helps reviewers curate faster and builds trust in the AI workflow. When you flag a test for deletion, cite the specific rule it violates.

How do we know it's better?

mbj/mutant repository on GitHub
https://github.com/mbj/mutant

HARNESS THE AGENT · MUTATION TESTING

Meet Mutant

Spawns tiny mutations in your code
Runs the test suite against each mutant
Inspect survivors — missing test or dead code?
Kill survivors with new tests or simpler code
Alex Mack morphing into liquid

HARNESS THE AGENT · MUTATION TESTING

Mutant in Action

lib/person.rb
def adult? @age >= 18 end
test/person_test.rb
test "adult at 19" do assert Person.new(age: 19).adult? end
test "not adult at 17" do refute Person.new(age: 17).adult? end
# ✅ Tests pass · 100% line coverage
lib/person.rb
def adult? @age >= 18 end
def adult? - @age >= 18 + @age > 18 end
test/person_test.rb
test "adult at 19" do assert Person.new(age: 19).adult? end
test "not adult at 17" do refute Person.new(age: 17).adult? end
# ✅ Tests still pass
Boundary case was never tested, the mutant survived

HARNESS THE AGENT · MUTATION TESTING

Mutant in Action

app/models/order.rb
def total_price if line_items.any? line_items.sum(:price) else 0 end end
test/models/order_test.rb
test "sums line item prices" do assert_equal 30, orders(:with_items).total_price end
test "empty order returns 0" do assert_equal 0, Order.new.total_price end
# ✅ Tests pass · 100% line coverage
app/models/order.rb
def total_price if line_items.any? line_items.sum(:price) else 0 end end
def total_price - if line_items.any? - line_items.sum(:price) - else - 0 - end + line_items.sum(:price) end
test/models/order_test.rb
test "sums line item prices" do assert_equal 30, orders(:with_items).total_price end
test "empty order returns 0" do assert_equal 0, Order.new.total_price end
# ✅ Tests still pass
The guard never changed the result, the mutant survived

HARNESS THE AGENT · MUTATION TESTING

Evil Mutant

Modified boundaries (>= → >)
Removed side effects
Changed return values
Deleted conditional branches
Add the Missing Test
Defensive code that can't trigger
Speculative branches nobody asked for
Redundant expressions (.to_s on a string)
Dead / unreachable code paths
Simplify the Code

HARNESS THE AGENT

Bolster the AI Harness

Deterministic Tools
Code scanning and coverage tools can objectively verify what the test suite actually catches.
Gemfile
source "https://rubygems.org" group :test do # the floor — "Did this code run during tests?" gem "simplecov" # the gate — "Is the code I just changed tested at all?" gem "undercover" # the deep check — "Would tests notice if this code changed?" gem "mutant" end group :development do # the sweep — "Is this code still alive?" gem "debride" # methods nothing calls gem "keela" # methods, scopes, constants, partials… with a baseline end group :production do # the cleanup — "Is this code even used in production?" gem "coverband" end

Fewer tests.More confidence.

thank_you_test.rb · passing
test "thanks for listening" do assert_equal :grateful, speaker.mood end # Finished in 0.09s 1 runs, 1 assertions, 0 failures, 0 errors, 0 skips ✓
iribon9@gmail.com linkedin.com/in/ifatribon
/ 20

Enter to jump · Esc to cancel

Keyboard Shortcuts

Next slide→ · Space · PgDn
Previous slide← · PgUp
First / last slideHome · End
Go to slideG
Presenter viewP
FullscreenF
Toggle this helpH · ?

Press any key to close