Oracle Exadata observability

Your Exadata May Be Healthy.Your Performance May Not Be.

What can be difficult to see when monitoring a highly consolidated Exadata environment.

Estimated reading time: 5–6 minutes

Why this matters

A healthy status does not tell the whole story

Exadata gives you a lot of information about what is happening in the environment. CPU, I/O, database activity, storage, network, and many other metrics are available.

So why can performance problems still be difficult to troubleshoot? One reason is that having the metrics is not the same as having the right view of the environment.

This becomes particularly noticeable on consolidated Exadata systems. When many databases share the same infrastructure, the performance of one database can be influenced by what another database is doing. A problem may also be short-lived and difficult to see in historical data.

This is where the limitations of a traditional monitoring approach can become visible.

You Have the Data. But Can You See the Story?

When investigating an Exadata performance problem, the first challenge is usually not finding a metric. It is putting the metrics together.

The database symptom

Suppose a database suddenly starts responding more slowly. Looking at the database itself may show increased latency or higher CPU usage, but that does not necessarily tell you why.

The platform context

At the same time, another database may be generating heavy I/O, a storage cell may be under increased load, network traffic may have changed, and several workloads may be competing for the same resources. All of these metrics may be available, but if they are viewed separately, the DBA has to connect the dots manually.

From predefined views to monitoring designed for Exadata

Oracle Enterprise Manager

Oracle Enterprise Manager provides a broad set of Oracle and Exadata monitoring capabilities, including predefined dashboards and views. For many operational tasks, this is sufficient. The problem appears when you need to look at the environment in a way that is specific to the incident you are investigating.

Grafana

Grafana provides much more flexibility for this type of analysis. Metrics can be combined, dashboards customized, and data from different sources displayed together. However, building this kind of Exadata monitoring environment yourself means deciding what to collect, how to collect it, how to retain it, and how to build and maintain useful dashboards.

GoMon4Exa

This is one of the areas where GoMon4Exa can be used as an alternative. It uses Prometheus for metric collection and storage and Grafana for visualization, but the dashboards and monitoring approach are built specifically around Exadata.

Instead of starting with a generic monitoring platform and adapting it to Exadata, start with the Exadata troubleshooting questions and build the monitoring view around them.

What If the Performance Problem Has Already Disappeared?

Some of the most difficult performance problems are the ones that are no longer happening when you start investigating them.

A storage or CPU problem can last a few minutes. An application may experience degraded performance, users report the issue, and by the time the DBA starts looking at the environment, everything appears normal again. This makes historical data extremely important.

OEM provides historical metric data, but detailed information is progressively aggregated over time. Collection intervals also depend on the metric and configuration. This is useful for understanding long-term trends. It can be less useful when the question is: what exactly happened during those five minutes?

A short CPU saturation, storage disturbance, or network issue can have a significant impact without making much difference to a daily average. This is where higher-resolution metrics become useful.

With Prometheus, metrics can be retained at their original collection resolution for longer periods. Instead of looking only at averages, you can investigate the actual sequence of events.

Five minutes that explain the incident

  1. CPU starts increasing

  2. database workload increases

  3. I/O latency changes

  4. application response time increases

  5. workload returns to normal

That sequence can tell you much more than a daily average. Prometheus and Grafana can provide this capability, but there is work involved in building an Exadata-specific monitoring setup.

GoMon4Exa provides this layer out of the box for Exadata, with Prometheus-based high-resolution historical metrics and Grafana dashboards designed for investigating Exadata performance.

The important point is not simply keeping more data. It is being able to reconstruct what happened after the problem is gone.

Your Database May Not Be the Problem

This is probably one of the most common issues in a consolidated Exadata environment. A database is slow, so naturally the first place to look is the database.

But the database may simply be where the problem becomes visible. If several databases share the same infrastructure, one workload can affect another through shared CPU, I/O, network, or storage resources. A database may experience higher latency because another workload is consuming significant resources on the same platform.

OEM provides tools that can help identify database resource consumption and understand how workloads use Exadata resources, particularly for CPU, I/O, and other Oracle-specific metrics. But in a highly consolidated environment, another question becomes important: who else was using the same resources?

Understanding individual workload contribution can support capacity planning and workload optimization and, where appropriate, provide the basis for usage reporting or chargeback.

GoMon4Exa was designed with this type of consolidated Exadata environment in mind. Rather than looking only at individual database performance, it provides a platform-level view that helps identify resource consumers and potential contention between workloads.

Starting question

“Why is this database slow?”

Better starting point

“What is happening on the platform that could be making this database slow?”

Exadata Doesn't Run in a Vacuum

Exadata is only one part of most production environments.

The source of the problem

An application performance problem can involve the application itself, Kubernetes, the host, the database, storage, network infrastructure, or another workload. This means that the root cause is not always visible in the Oracle monitoring environment.

A shared timeline

Prometheus and Grafana are useful in this type of architecture because they can bring metrics from different systems into one monitoring layer. You could put application response time, database activity, host CPU, storage I/O, and Exadata metrics on the same timeline.

Platform configuration

A generic observability platform still needs to be configured for the specific environment. You need the right exporters, metrics, labels, retention, and dashboards. You also need to know which Exadata metrics are actually useful when investigating a particular problem.

The Exadata layer

GoMon4Exa provides the Exadata-specific part of this model. Because it is based on Prometheus, Exadata metrics can also be used alongside metrics from other systems rather than creating another isolated monitoring environment.

Did these two things happen at the same time, or are we just assuming they are related?

The aim is not to replace every monitoring tool. It is to make Exadata a useful part of the same observability picture.

Finding the Slow Database Isn't Finding the Root Cause

When troubleshooting, there are usually two directions from which you can approach the problem: from the platform or from the workload.

Top-down

Start from the platform

You see that something is wrong with the Exadata environment. You look for unusual resource consumption, identify the main actors, and then drill down to the database, workload, and eventually the SQL responsible.

Bottom-up

Start from the workload

You already know which database or query is experiencing a problem. You ask what resources it is consuming, what other workloads use those resources, what is happening on the storage or host, and which other databases could be affected.

Both are useful, because the source of the problem and the visible symptom are not always in the same place. A database may be reporting high latency while another workload is responsible for the resource pressure.

This is where a monitoring approach focused primarily on individual database performance can become limiting. GoMon4Exa supports both directions of investigation, allowing you to move from the Exadata platform toward individual workloads and SQL, or start with a specific database and investigate its impact on the wider platform.

For a DBA, the practical benefit is simple: less time moving between unrelated views and more time understanding what is actually happening.

From monitoring data to an explanation

Monitoring should explain, not only alert

Exadata already provides a lot of monitoring data. The challenge is making that data useful when something goes wrong.

Oracle Enterprise Manager is strong when it comes to Oracle-specific monitoring, diagnostics, and management. Prometheus and Grafana provide flexibility, high-resolution metrics, and the ability to combine data from different systems. Neither approach necessarily has to replace the other.

For organizations that need a more Exadata-focused observability layer, GoMon4Exa provides another option, combining Prometheus-based metrics with Grafana dashboards designed around Exadata performance analysis and consolidated environments.

“Something is wrong.”

“We know what happened, why it happened, and what it affected.”

That is the difference between monitoring Exadata and actually being able to troubleshoot it.