Process Mining for Automation Candidate Selection: What the Event Log Should Disqualify

Process mining can reveal whether an automation candidate is genuinely repeatable, observable and stable. Learn how event logs expose variants, rework, weak data and fragmented paths before engineering begins.

Introduction

A process with high volume, long cycle time and obvious manual effort can still be the wrong automation candidate if its event log shows no stable path worth encoding.

Automation discovery often begins with workshops, employee interviews or lists of repetitive tasks, where a process can appear consistent from the perspective of one team even though actual cases travel through several materially different routes.

Process mining changes the evidence available for that decision. Microsoft’s Process Mining documentation describes capabilities for comparing processes, identifying root causes, detecting inefficiencies and finding opportunities for improvement and automation. SAP Signavio’s execution-variant analysis similarly exposes the different sequences cases follow and their effect on process performance.

For process mining automation candidate selection, that evidence should be allowed to eliminate candidates that are too fragmented, poorly observed or operationally inconsistent to justify development in their current form—not simply populate an automation backlog.

At a glance

  • Process names and aggregate transaction volumes can conceal materially different execution paths.
  • Variant distribution reveals how much genuinely repeatable work exists inside the process.
  • Frequency becomes useful only after the automatable activity or path has been isolated.
  • Rework and conformance analysis can reduce the scope of a candidate before engineering starts.
  • Incomplete event data can make the mining result itself too weak to support an automation decision.

Automation candidate selection starts below the process name

Business processes are convenient units for ownership and reporting, but they are often too broad for automation selection.

“Invoice processing” can contain straight-through invoices, invoices without purchase orders, quantity mismatches, blocked payments, credit notes and supplier-master problems. The cases share an organisational label without necessarily sharing one executable route.

SAP Signavio’s execution-analysis documentation uses supplier invoicing to show how cases can follow different paths depending on whether invoices are created manually, through EDI/IDoc or through a BAPI interface. Other variants can include payment blocks and manual releases. Treating those event sequences as distinct execution variants exposes a distribution that a single process name would hide.

A process with 100,000 annual cases may contain a 60,000-case stable path, several smaller paths requiring different controls and a long tail of exceptional cases. The relevant denominator for automation is the population that follows the specific sequence the proposed automation can actually execute.

The automatable unit may therefore be one path or recurring activity rather than the entire process.

The dominant path matters more than the process label

Variation does not automatically make a process unsuitable for automation; its distribution determines what kind of automation boundary is realistic.

A process in which three or four variants account for most cases presents a different engineering problem from one in which activity is dispersed across dozens of materially different paths.

Current SAP Signavio Process Intelligence capabilities allow teams to explore variants and compare process behaviour using measures such as cycle time, conformance and other case-level metrics. Its process metrics documentation also includes measures covering automation potential, changes, cycle time and conformance.

That gives automation teams evidence about where standardisation actually exists rather than requiring them to infer it from a procedure manual or stakeholder description.

A dominant path may justify automation even while the wider process remains heterogeneous. Conversely, a process that looks standard in documentation can prove highly fragmented once actual execution paths are counted.

The candidate boundary should follow the repeated behaviour visible in the data, not the label attached to the process.

Volume means little until the unit of work is isolated

High transaction volume is one of the strongest arguments routinely made for automation.

It becomes useful only after teams know what is being counted.

A study of a loan-application process provides a useful example. Researchers began with an event log containing 24 activities and assessed RPA suitability using manual work, software involvement, frequency, productivity, rework and process-flow characteristics.

Mandatory criteria reduced the field to seven activities. Frequency then removed two more: one had occurred only 108 times and another only twice. The remaining candidates were assessed using additional operating evidence rather than treating transaction count as sufficient justification. The full method is described in the loan-process RPA assessment study.

The sequence matters because frequency belongs to a specific activity or path.

A task performed 40,000 times across several applications, rule sets and user groups may be less attractive than a smaller population with highly consistent inputs and actions.

Process mining helps establish what the volume figure actually represents before it enters an automation business case.

Rework can shrink the automation boundary

Repeated work can look attractive because it represents labour that might be removed, yet the event log may show that the repetition is downstream of another problem.

A case returns to validation because required information is missing. An approval repeats after a downstream change invalidates an earlier decision. Data is re-entered because another system rejected the first transaction.

Microsoft’s 2026 release-wave documentation for Power Automate process mining lists capabilities including rework detection, root-cause analysis and process comparison.

Some repetitive correction work is itself a viable automation candidate: predictable, frequent and governed by clear rules. In other cases, the repeated activity sits downstream of a condition the automation would leave untouched. Automating that step may reduce handling effort while preserving the mechanism that keeps sending cases back to it.

Process mining cannot decide which business condition should change. It can show where the recurrence appears, which cases experience it and what tends to happen beforehand.

That evidence may turn a proposed end-to-end automation into a much smaller candidate.

Conformance can isolate the part worth automating

A procedure describes the intended process.

The event log records what occurred.

The difference becomes useful when automation teams need to determine whether deviations are legitimate operating variants or signs that one automation would have to absorb too many unrelated conditions.

SAP Signavio supports conformance analysis through process metrics and analysis configuration. Its current analysis-configuration documentation allows conformant and non-conformant behaviour to be identified and used in analysis, while its standard metrics include conformance level and case counts associated with compliance risk.

Suppose 80% of cases follow a consistent, compliant route and the remainder split across several deviations requiring investigation.

The organisation does not have to choose between automating the entire process and rejecting it. The stable population can be assessed as one candidate while the remaining cases stay outside the automated boundary.

That is often a better engineering decision than translating every observed deviation into another branch of the same automation.

An event log can make a process look cleaner than the work really is

Process-mining output carries the authority of operational data.

Its coverage still has limits.

Microsoft’s guidance on preparing data for process mining requires case identifiers, activity information and timestamps to reconstruct process instances. The same documentation explicitly warns that application databases may contain only current state rather than historical events and that not every event of interest is necessarily logged.

That qualification can materially change candidate selection.

A procurement platform may record when a request was submitted and approved without recording the email exchange used to resolve missing information. A CRM may capture status transitions but omit spreadsheet work between them. A case-management system can record a final decision while the research supporting it takes place elsewhere.

The mined path may consequently appear more deterministic than the real work.

A dominant system variant is weak evidence for automation if the unobserved portion contains the judgement, data repair or cross-system activity that determines whether a case can proceed.

Where that gap is material, task-level evidence may be needed. Microsoft’s process- and task-mining guidance distinguishes organisation-level event-log analysis from task mining that captures desktop activity and user actions.

Candidate confidence should reflect how much of the work is actually visible.

A credible shortlist gets smaller as the evidence improves

Automation discovery often carries an implicit expectation that analysis should produce more opportunities, which can bias the exercise toward addition rather than elimination.

A better discovery process removes opportunities that looked attractive before the operating evidence was examined.

Academic work supports that approach. A systematic review of 32 studies on process mining and RPA found that process mining can strengthen automation discovery while also highlighting recurring problems with event-data collection, preprocessing and detailed routine discovery.

The evidence may therefore produce several legitimate outcomes. A candidate can proceed unchanged, its scope may be narrowed to one variant or activity, more data may be needed before the decision is credible, or the process may be removed from the backlog.

Rejecting a candidate after analysis protects engineering capacity from work whose repeatability or observability was never established.

Screen the candidate before calculating ROI

Automation programmes often move rapidly from discovery into prioritisation.

Candidates receive estimates for volume, time saved, implementation effort and expected return.

Those calculations become misleading when the unit being scored is still unstable.

A high-volume process may contain little repeated work at the level the automation would actually perform. A process with long cycle time may spend most of that time waiting. A heavily manual workflow may depend on activities that never appear in the event log.

The loan-application study avoided this problem by reducing the candidate population before deeper prioritisation. Activities that failed mandatory suitability characteristics were removed first, followed by candidates with insufficient frequency. Only the surviving set justified further assessment.

ROI analysis becomes more meaningful once the organisation has established that a coherent automation candidate exists.

The automatable unit is often smaller than the process

Process mining can produce a result that looks less ambitious and is technically more useful.

Instead of automating “invoice processing,” the evidence may isolate one recurrent invoice path.

Instead of automating “loan applications,” it may expose two repeatable activities inside a much larger case journey.

Instead of trying to accommodate every customer-service variant, it may reveal one population whose inputs, rules and system actions are consistent enough to justify automation.

This is also where candidate selection should stop.

Whether that unit of work should ultimately use RPA, a workflow engine, an API or another automation technology is a separate architecture decision. The distinction between those responsibilities is explored in One Enterprise Process Can Need RPA, a Workflow Engine and an AI Agent.

Candidate selection has a narrower job: establish that the work is sufficiently observable, repeated and bounded to deserve technical design.

Wil van der Aalst makes the broader point in his paper on hybrid intelligence and process mining: process mining is useful for deciding both what to automate and what not to automate as work is redistributed between people and software.

The strongest discovery programme is not the one that generates the longest backlog.

It is the one that sends fewer weak candidates into engineering.

FAQs: process mining and automation candidate selection

How does process mining help identify automation candidates?

Process mining reconstructs observed process behaviour from event data.

Teams can use that evidence to examine variants, frequency, cycle time, rework, conformance and other characteristics before deciding which activities or paths deserve automation assessment.

Does high transaction volume make a process a good automation candidate?

Not by itself.

Volume becomes meaningful after teams establish that a sufficiently large population follows repeatable rules and a consistent execution path.

Should a process with many variants be rejected?

Not automatically.

Variant analysis may show that a small number of paths cover most cases. Those paths can be assessed separately rather than forcing every variant into one automation.

How does rework affect automation selection?

Rework can expose repetitive work worth automating, but it can also show that cases repeatedly return to a step because of an unresolved upstream condition.

The pattern should be understood before the repeated activity becomes an automation candidate.

What data does process mining require?

Case-based process mining normally requires a case identifier, activity information and timestamps.

Additional attributes add business context, but the result can remain incomplete if important manual or cross-system activity is not captured in the event log.

Can process mining decide automatically whether a process should be automated?

No.

It provides evidence about observed process behaviour. Business value, control requirements, technical feasibility, planned changes and human judgement still influence the final decision.

Good process mining should shorten the automation backlog

An automation discovery exercise that identifies dozens of opportunities can look productive.

A more valuable exercise may conclude that one candidate is too fragmented, another relies on incomplete event data and a third contains too little repeatable activity to justify development.

A fourth has a stable path, sufficient volume and enough observable evidence to warrant engineering.

That one moves forward.

Process mining has done useful work before a single automation is built.

Automation candidate selection improves when the event log is allowed to remove weak work from the pipeline before development begins.

Leave a Reply

Your email address will not be published. Required fields are marked *