# Common DevOps Pain Points

> Common IT problems solved with DevOps: downtime, rollbacks, test systems, infrastructure as code, and canary releases.

Source: https://www.xalt.de/en/blog/common-devops-pain-points/

TEAM XALT Atlassian Platinum Partner · 15 June 2021 · 10 min

DevOps – a buzzword that has recently become a talking point in many companies. The spotlight of the IT industry has shifted to DevOps in recent years from “corporate disruption” and “agile everything”. How can the DevOps vision be reconciled with the reality of companies.

DevOps is the abbreviation for a set of methods, tools and processes for the automated deployment of software or services for a company. DevOps can support services over a longer period of time, considering the entire life of the software. These include, for example, installation, configuration, error handling and also deployment for users. Topics such as rapid system availability, adherence to update cycles and also the provision of new infrastructure fall into this category.

## **Automatic Load Monitoring for Atlassian Plugin Updates**

Although most Atlassian Marketplace plugins are thoroughly tested for stability, performance and load, it happens again and again that Jira and Confluence instances require additional resources during an update.

Especially if you carry out custom developments in your instance, this visibility is indispensable.

Additional load becomes problematic when it slows down the instance and has a noticeable impact on user experience and system stability. In the worst case, this leads to a complete failure of production systems and downtimes.

With the right tools and strategies from the methodology summarised under DevOps, this belongs to the past:

- With continuous monitoring and data collection, you can directly see the impact of updates on your instance
- React to performance drops by being able to roll back updates.
- With Blue-Green Deployments and DevOps, your system can additionally be tested for performance by, for example, gradually letting users onto the more current system (This is known as Canary Releases - you will learn more about this below).

## **Rapid recovery during downtimes**

At leading Fortune 1000 companies, the average total cost of unplanned system outages per year ranges from $1.25 billion to $2.5 billion. This means that every hour a system is unavailable costs $100,000.

These costs are not only annoying but can usually be easily avoided. DevOps provides a framework that uncomplicatedly minimises possible downtimes and uses technologies to get back online in the shortest possible time in the event of a system failure.

DevOps helps to minimise downtimes and accelerate recovery by leveraging Spinnaker technology at XALT. This makes it possible to quickly and easily roll back to previous versions of the system without losing important data in the database.

## **System availability within minutes**

In software development, various systems are often needed for development or testing without prior warning. These systems or IT infrastructure are normally configured manually without DevOps. However, manual configurations usually require a lot of time and are more prone to errors.

Through automated infrastructure provisioning (with IaC and DevOps), all parts of your IT infrastructure (e.g., servers, databases, load balancers, and systems) are no longer manually configured or managed by developers in the cloud provider's GUI every time.

Benefits of Infrastructure as Code:

- Lower costs
- Faster deployments
- Fewer errors
- Better infrastructure consistency
- No configuration drift
- Examples of IaC tools (Ansible, Terraform, CloudFormation, Helm)
- IaC not only automates the process but also serves as a form of documentation for the correct procedure when deploying infrastructure

For IaC, existing tools for server automation and configuration management can often be used. Additionally, there are solutions specifically developed for IaC.

## **Availability of development systems** and test systems

Developers constantly test new features, updates, and bug fixes on development and test systems. And what impact or effects these “changes” have on the system.

However, these “changes” often need to be compared with one another to draw conclusions for recommendations on action. With the help of IaC, an unlimited number of development systems can be generated and provisioned. For example, it is easy to deploy another Kubernetes cluster or, using Helm, another application.

By having several different development systems active simultaneously, your IT team can always draw comparisons between the individual systems without your production system (live system) being affected.

Through the use of “Infrastructure as Code”, new systems can also be spun up at any time within moments for use in development.

Furthermore, systems that are no longer needed can be shut down at any time as soon as they are no longer required.

Advantages of multiple development systems

- Real-time comparison of multiple, yet internally different, development systems.
- Derivation of recommendations on action through analysis of comparison data.
- Direct availability of additional development systems on demand through the “Infrastructure as Code approach”.
- Each developer or team has their own test system

## **Development systems are not comparable to the live system**

Many developers face the problem that the system they use for development is often not comparable to the live system. Important parameters, plugins, and integrations are often missing, or only an old dataset is available.

To ensure that your developers always have access to the current dataset, or a comparable system, we rely on "Infrastructure as Code". This makes it possible to provision a system for development ad hoc. On the other hand, the live system is cloned and the current dataset is used.

![Continuous Integration & Delivery and Blue-Green Deployment](https://cdn.sanity.io/images/c475o02b/production/6f1ada8721f475faa7642bfb3a5787c53a35e08a-1600x1200.png?w=1504&q=75&fit=max&auto=format)

Green aktiv nach neuem Release

Subsequently, your developers can add further features to the cloned system or update it. This is where the Blue-Green approach comes into play again. The cloned system is then merged with the Green system. When the time comes, the production system is switched from "Blue" to "Green".

The advantage of this approach lies primarily in the fact that your developers can not only always work on current and comparable systems, but the development systems can subsequently be put live. This saves your company not only a lot of time but also a considerable amount of resources and accelerates your processes.

## **The Staging System corresponds to the Development System**

Before a system goes live, it is usually presented for acceptance via a staging environment. This typically gives the responsible party (or customer) the opportunity to verify the functionality of the new system, updates, or feature integrations. In many cases, these staging environments are set up specifically for this purpose and are only needed for that.

By using DevOps and the ability to provide systems ad hoc, this problem can be easily eliminated. Additionally, a development system can also be used as a staging system/environment and subsequently released as an active and, above all, stable system via a so-called Blue-Green Deployment.

## **Simple rollback to a previous version with Blue-Green and Spinnaker**

Even if updates, code changes, or feature integrations have been tested multiple times, it still happens that the system does not function correctly in active operation, contains bugs, or does not perform as hoped. In this case, it is often only possible to perform a rollback to the previous version to ensure the stability of the system and have a functional system again. Subsequently, it can be investigated why the product system no longer functioned.

Without the use of Spinnaker or the Blue-Green approach, rollbacks require a lot of time and cost money in the form of additional downtimes.

Therefore, XALT uses the tool Spinnaker. Spinnaker is a multi-cloud continuous delivery platform for releasing software changes, which can store multiple versions of your systems.

This allows a previous version to be restored with just one click. This process is further accelerated by using the Blue-Green approach. In this case, one or more similar systems are always in circulation.

## **Automated testing of new versions of external software (e.g. Jira and Confluence)**

Your company probably uses some external software solutions that are installed on an AWS or Azure cloud instance. However, updates to these apps are not in your hands. And they can often lead to unwanted problems during updates. For example, plugins or integrations may no longer work. REST API changes lead to errors in individual connections to other tools. Or the update is simply not mature and leads to performance drops on your instance.

![CI/CD Workflow und Feature Branching als DevOps Pain](https://cdn.sanity.io/images/c475o02b/production/2f1c77d1ddaed5ceff610517d712ad2002c866a8-1920x1080.png?w=1504&q=75&fit=max&auto=format)

**Example: Automated testing of Confluence updates**

To manage the situation, a new branch with a copy of the main system is created before an update, and an update is performed. The live system is not affected by this. Subsequently, some tests are conducted in the new branch. Load tests to measure performance. Regression tests to find out if all active plugins are compatible with the new Confluence version.

## **Automatic maintenance page during downtime or unavailability**

You probably know it. You try to visit a website or web service and end up on a blank page without information. Is this due to my connection, or is the server unreachable?

Such situations are usually very unpleasant and often lead to numerous support requests to your service staff.

This problem can usually be solved with a simple maintenance page, which is automatically activated when the system is unavailable.

An effective approach to automatically switch to maintenance pages can look as follows:

1. An automatic **system health check** is carried out at regular intervals.
2. If the health check is successful, the system continues to run.
3. If the health check is not successful, the maintenance page is displayed and further information is provided to the user.
4. The IT department then carries out a comprehensive test of the system and makes it available to users as quickly as possible.

## **Automated testing of backups to ensure system stability**

Normally, the database is shut down for a few moments before each backup to prevent errors in the backup. Another problem that can arise is that data is lost when copying the database. If the system is subsequently restored from this backup, it is often not possible to start it up, which leads to further downtimes and further restricts availability.

To avoid this problem, automated backup tests are run to check integrity. A health check is performed and the backup is tested for completeness. If both of these tests are successful, the system is started up again and made available to users.

With this approach, we create a double safety measure for rolling back to a previous version and achieve higher availability of production systems.

## **Canary releases - gradual system rollout for all users**

Especially in systems with thousands or tens of thousands of simultaneous users, updates or new features can lead to poor performance or availability. A widespread problem is that new features work smoothly in test environments but are only usable to a limited extent under full "load" or in the live system. New features often have the characteristic of being quickly wanted to be used by a large number of users. In doing so, it can happen that the server or app cannot handle this surge. In the worst case, the system then crashes.

However, this problem can be solved using so-called canary releases. With the help of several parallel running systems (blue-green), these features are gradually enabled for users. This allows it to be checked whether the system can withstand the load and maintains its integrity and stability. The advantage is that it is possible to return to the previous version at any time and your users will not notice any impact on availability.

## **Confluence: read-only version during downtime, update or bug fix**

#### **Intranet and wiki always available**

It is always extremely frustrating when important systems are temporarily unavailable. This is true for users on the one hand, and for your developers and Confluence / system admins on the other. Your users then no longer have the opportunity to look up important information from your wiki, collaboration with your team is restricted, and in the worst case, missing information has a negative impact on your business success.

However, at XALT we have found that most Confluence users "only" read in the company wiki, but do not write their own contributions or edit content.

The question that arises from this is, therefore, how do I provide all my employees with all the relevant information if Confluence goes down.

Confluence offers a so-called "Read Only" version for this purpose. This is normally used when your admins activate maintenance mode. The great advantage is that the entire wiki, including the database, remains available to all users at all times.

However, the Read Only mode must first always be activated manually, which is difficult to achieve when the system is inactive.

Using the DevOps methodology, we can support you in providing a Read Only version automatically in the event of downtimes. This allows your employees to continue accessing the system and obtaining important information at any time.

## **Automated update rollout at consistent time intervals**

In many companies, updates are manually installed on active systems, and time slots for updates are predetermined in advance.

Many updates run in the background and do not affect system availability or only for brief moments. However, some updates require the system to be shut down for several hours.

By implementing automation and using DevOps tools, these manual steps become a thing of the past.

At XALT, we use the Blue-Green deployment methodology and approach to ensure 99.95% availability. And we also use a process that merges finished development systems with the inactive "Green" instance and releases them at a fixed time, meaning it switches from Blue to Green.

This not only accelerates the time from development to deployment and provides your employees with a real time advantage, but also ensures that the former live system (Blue) offers the possibility to perform a rollback quickly and easily.

Do you recognize some of these problems in your own organization and want to get a handle on the situation? Our experienced DevOps engineers are happy to help you and provide initial recommendations in a first conversation on how you can integrate DevOps early on.

Further information on DevOps can be found here: [DevOps Transformation](https://www.xalt.de/en/devops-transformation-implementation/)

## Recognize these problems in your own setup?

Our experienced DevOps engineers can give you a first set of recommendations for adopting DevOps early.

[Talk to us](https://www.xalt.de/en/contact/)

## More articles

### [The Agent Harness: The Real Foundation of Enterprise AI](https://www.xalt.de/en/blog/agent-harness-foundation-enterprise-ai/)

Why the agent harness – not the AI model or the generated code – determines the success of enterprise AI initiatives, and how to build one in five steps.

*DevOps*

### [Governing Agentic AI Safely: Zero Trust and Compliance for AI Agents](https://www.xalt.de/en/blog/governing-agentic-ai-zero-trust-compliance/)

AI agents are transforming the enterprise at speed – and the risk is growing just as fast. How the "in dubio pro securitate" principle, formal verification, and zero trust keep agentic AI safe and compliant.

*DevOps*

### [Shift Left: Catching Bugs Earlier in Your Development Cycle](https://www.xalt.de/en/blog/shift-left/)

Bugs found late cost time, money, and trust. The shift-left approach moves testing and quality assurance as early as possible into the development cycle.

*DevOps*

---

*This version is for AI agents. Every page of this site is available as Markdown: append `index.md` to its path. Index of all pages: [llms.txt](https://www.xalt.de/llms.txt)*
