Krzysztof Kaszanek from Kiwee alongside a sloth illustration and geometric mountain outlines.

The Real Cost of Not Testing Your E-Commerce Shop

There is a conversation we at Kiwee often have with eCommerce teams, and it usually starts the same way. The shop works, plugins are running, orders are going through. And when something needs checking before a release, a developer clicks through the checkout flow, it looks fine, and the update goes live. So why would you invest time in automated testing?

Manual testing is not free

It's a fair question, but most teams frame it wrong: as an added cost. Manual testing isn't free, it just feels that way, because the cost shows up later instead of on an invoice. Ship a plugin or a custom integration without automated tests, and you're not avoiding that cost, you're just paying it differently. Something always changes eventually — a Shopware update, a shifted dependency, a payment provider tweaking their API — and someone has to click through the same flows again to check nothing broke. That's real time, and time is usually the most expensive part of any project. With AI now producing more code than any team could write by hand, that manual-checking cost only grows.

Chart showing manual testing costs increasing over time, while automated testing costs decline after a higher initial investment.

What a release looks like without CI/CD

For a lot of eCommerce shops, a major platform update looks something like this: you block out half a day, make a backup, mirror the old files somewhere in case you need to roll back. You apply the update, run through a checklist manually, disable maintenance mode, do another round of testing, and then hope nothing slipped through.

Every release is a multi-step manual process that carries real risk.

With a proper CI/CD pipeline in place, that manual complexity disappears and so does most of the room for error. You push a tag, deployment is triggered, maintenance mode is automatically enabled, and you're done. Half a day shrinks to about half an hour, with maybe ten minutes of real downtime for customers. And the bigger win is consistency — the process runs the same way every time, so there's no guessing what might go wrong from one release to the next.

Not only is that a developer quality-of-life improvement, but, more importantly, it's increasing the operational reliability of the business. When a release stops being a stressful event and becomes routine, teams ship more often with confidence.

The three layers of testing

Automated test structure is often represented visually as a pyramid, consisting of the following layers:

  • Unit tests cover small, isolated pieces of logic. They're fast, you can run hundreds of them in seconds, and they form the foundation of a solid test suite. If your shop converts order data into a format your ERP expects, unit tests make sure that conversion keeps working correctly even as the underlying platform changes around it.

  • Integration tests check whether different parts of the system work together — whether a discount code applied at checkout actually reduces the price correctly, whether a payment confirmation from the gateway updates the order status, whether stock levels sync back to the shop after a sale.

  • End-to-end tests simulate a real customer session in the browser — searching for a product, adding it to the cart, going through checkout, completing the order. These are the most expensive tests to maintain, so you keep them focused on the critical paths, the ones where a failure directly costs revenue.

A healthy test suite has a lot of unit tests, a reasonable number of integration tests, and a smaller, deliberate set of end-to-end tests covering the critical flows that cannot break.

Testing pyramid showing unit tests at the base, integration tests in the middle, and end-to-end tests for critical customer flows at the top.

Where to start if you have an existing project with no tests

Looking at a large, mature codebase carrying years of technical debt, it's hard to know where to start. You might think you need to test every piece of code before it's safe to move forward. But that is not how it works in practice, and it is not the best use of your time.

For an existing project, it's best to start at the top. End-to-end tests on your most critical flows first.

  • Can customers create an account and log back in?
  • Can customers find and search for products?
  • Can they add the products to cart and complete checkout?
  • Do orders reach your ERP correctly?

Get those covered. Then work downward into integration tests for your external connections, and unit tests for the logic those connections depend on. Make sure all new code is tested from the start — that way coverage builds naturally over time.

For new features specifically, test-driven development (TDD) is worth considering. Writing the test before writing the code forces you to think clearly about what the code is actually supposed to do. The tests act as a specification the code has to meet. This approach becomes especially valuable in the times of AI-assisted development, which we'll get into in the next paragraphs.

How AI changed the testing landscape

The rise of AI coding agents has fundamentally changed the nature of technical debt. It's no longer just about slow, human-written code — it's about "machine-speed" debt. AI agents can generate features and implement changes at an unprecedented velocity, which sounds like a dream for productivity. However, these agents are also prone to generating insecure patterns, or code that is logically consistent but architecturally unsound.

If your testing strategy is still based on manual verification, you are now the bottleneck. Without a rigorous, automated test suite, you are essentially letting a machine push unverified code into production at a rate no human team could match. In an AI-first workflow, when up to 100% of code is generated by AI, your automated tests serve as the critical guardrail, ensuring that the speed gained by using agents doesn't come at the cost of stability.

Chart showing AI-generated code volume rising rapidly while human review capacity stays flat, creating a testing bottleneck.

That guardrail only holds if the tests themselves are trustworthy. The problem is that the tests are often AI-generated too. AI-generated tests are good at looking like coverage without actually being coverage. Ask a model to write tests for a function, and it will often test the happy path well and call it done. Worse, it can end up testing what the code does instead of what it's supposed to do. If the code has a bug, the test can lock that bug in by asserting the buggy output is correct. A green test suite built this way just proves the model patted itself on the back — the code agrees with itself, not that it's actually right.

Adjusting to the new reality

None of this means you should skip AI-generated tests. AI is genuinely good at scaffolding test files, coming up with edge cases a developer might miss, taking care of the boilerplate that makes developers avoid writing tests in the first place. We just need to pay even more attention to the testing and review discipline.

With the amount of code generated by AI, and the cognitive overload that comes with it, it's impossible to review code with the same level of detail as before. It doesn't help that AI is good at writing code that looks right — clean formatting, reasonable variable names, the kind of structure you'd expect from a good developer — which can give a reviewer a false sense of security. So tests are one area that might be worth most of our focus when reviewing the changes. We often say that tests are the best documentation, and it becomes even more relevant now. If you don't read through the generated tests carefully enough, you can't really be sure what is being shipped.

This is also why spec-driven development is gaining traction. Having a concrete feature specification was always important, but we know the reality — projects where the features are not really well defined, where teams rely on unspoken rules that "everyone just knows". Unfortunately, that won't work with AI. Unlike humans, AI has no knowledge of the business context, decisions that were made in the past or the reasoning behind a customer-specific exception. That's why it's crucial to put more effort than ever on creating detailed and well-structured requirements, even if some of them seem obvious. If you do that well, you will have E2E test scenarios for free, and AI can take the burden of turning them into actual test code.

AI-generated code quality

There are other quality risks too. Code style that's inconsistent from one file to another, overcomplicated solutions, patterns pulled from training data that doesn't reflect what is used in your project or dependencies that don't quite match the rest of the project.

None of these are new problems, they're the same issues any team with loose conventions has always had. AI just makes them appear faster. The fixes are familiar too:

  • Linters and formatters that catch style drift automatically.
  • Static analysis and type checking that catches the bugs the "looks correct" code likes to hide.
  • Dependency and security scanning, to catch known vulnerabilities in whatever version AI happens to suggest.

This is the idea behind what's being called maintainability sensors: none of the tools themselves are new, what's changing is how we need to integrate them into our workflow. Instead of running them on commit, push, or pull request creation, we should build them into the AI's coding session, giving it a feedback loop it can immediately act on while working.

Sensors can't catch everything, though. For example, there is no sensor for overcomplicated solutions — that depends on a human review that actually reads what's being implemented, not just that the checks passed. No amount of tooling replaces a reviewer who knows the codebase well enough to recognize when a simple problem got an overly elaborate solution.

When a bug appears anyway

Testing doesn't eliminate bugs. What it changes is how you respond to them, and how many of them come back.

When a bug is found, the right response is two things: fix it, and write a test that covers that specific case — including the edge case that caused it, not just the happy path where everything behaves as expected.

Every bug that gets a test attached becomes a bug that cannot quietly reappear months later when someone updates a dependency. Over time this builds up. The test suite grows not just with new features, but with every problem you have resolved. The codebase becomes progressively harder to accidentally break, not because it is perfect, but because there is an expanding safety net underneath it.

Is your release process the bottleneck?

There's a benefit to testing that doesn't show up neatly in time-saved calculations, but it's probably the most important one in practice.

Without tests, releases run on a kind of quiet hope. You push the code and hope nothing broke. You go through the checklist and hope you didn't miss anything. Tests replace that with proof — evidence that the code actually does what it's supposed to do. You can run them before a release, after a dependency update, after a refactor, and see whether the behavior you expected is the behavior you have. That proof builds confidence and lets a team move at the pace the business actually needs.

Dashboard showing unit, integration, end-to-end, visual, and performance tests passed, with a release-ready status and Deploy button.

Most teams without that confidence don't think of themselves as teams that don't test. They think of themselves as teams that test when it matters, manually, before releases, carefully. And that might work for a while.

The moment it stops working is usually not dramatic, it's gradual: releases start taking longer, updates get delayed because nobody wants to be the one who broke checkout, and everyone gets a little more nervous about touching certain parts of the codebase. AI-assisted development can get a team here faster than ever, because more code moves through the pipeline.

If any of that sounds familiar, the problem is usually not a specific bug or a specific gap in the team. It's the release process itself, and the lack of a safety net underneath it.

This is a large part of what we work on at Kiwee: helping eCommerce teams build deployment pipelines that are stable and reproducible, setting up staging environments that actually reflect production, and putting automated tests in place for the flows that matter most. Not as a one-time project, but as part of building a shop that can actually evolve without everything feeling fragile.

If you are heading into a Shopware migration, dealing with integrations that feel risky to touch, or just want to understand where your current setup is most vulnerable, we are happy to take a look.

FacebookTwitterPinterest

Krzysiek Kaszanek

Lead Software Engineer

I'm a Lead Software Engineer at Kiwee. I got into programming more or less by accident, after finding a C++ tutorial online. I enjoy solving problems, and this turned out to be a great way to do it.

Over the years I've worked on many different sides of eCommerce, and mostly specialize in Shopware. What I enjoy most is working with clients and turning a rough set of requirements into a real, working feature. I also like trying new things, especially on real projects.

Outside of work, I spend a lot of time outdoors. Rock climbing is my favorite way to relax.