Data Harmonisation with AIoT: What Do AI Agents Need for the Mapping?

Episode 224Sep 23, 202647 min

Madeleine Mickeleit talks to Dr. Jürgen Krämer, CPO and Managing Director at Cumulocity, on the IoT Use Case Podcast about AIoT: classic machine learning, generative AI and agentic AI meeting industrial IoT. The focus is the path from connecting machines to partly autonomous operations.

Summary

For Krämer, a securely connected fleet is a precondition, not an achievement: without it, no AI project is worth starting. He separates two kinds of AI: classic machine learning for high-frequency machine data, where an LLM would blow up token costs, and agentic AI, which starts a root cause analysis when an anomaly appears. Both stand or fall on context: without asset relationships, units and documentation, a fault code like 4711 on a pump means nothing to an agent.

Data harmonisation makes this tangible. In one plant the temperature sensor is called Hugo, in the next Willi. An Expert Agent reads the incoming data stream and proposes how to map it into the Cumulocity format – tested with OPC UA, Modbus, LoRaWAN, Protobuf, CBOR and Sparkplug. At Enercon, a rule-based remote action automates around 40 percent of turbine restarts; by Krämer's own account, no AI is involved.

Autonomous intervention remains future work: prototypes yes, production no. Customers bring their own LLM, and Cumulocity trains no models on customer data, protecting the operators' domain knowledge. Krämer also warns against vibe coding in business-critical processes: AI-assisted development belongs at the application level, not in the system of record that carries governance and auditing.

What you'll take away

  • No AI project pays off without a securely connected fleet: bad data produces bad results.
  • High-frequency machine data belongs with classic machine learning, not an LLM – token costs would spiral.
  • The semantic layer is industry-specific, which is where domain partners come in.
  • Data sovereignty means asking whether the vendor trains models on your operating data.

Transcript

Madeleine:

Hello, IoT friends. Welcome to the podcast about putting IoT to work in practice. I'm your host, Madeleine Mickeleit, a mechanical engineer by training. And today we finally get to a question a lot of you have asked me: AI versus IoT – do they belong together? And where exactly is the technical line between them?

There's also a new buzzword around that some of you will have heard and others won't: AIoT. Today we'll define it for you, and we'll answer a few questions from the community that have come my way. And as always, practice comes first – we'll explain it through what other companies are actually doing: what is already up and running, and what is still some way off.

With me today is someone I'd call a veteran of the IoT world – I hope he'll take that as a compliment: Dr. Jürgen Krämer. He's an entrepreneur, CPO and Managing Director at Cumulocity. Many of you will know the company from its Software AG days – we'll come to that in a moment. So today we're looking at IoT software: where AI is already finding its way in, which use cases exist, and which ones you should definitely know about. Let's get into it. As always, you'll find everything about projects like this at www.iotusecase.com – and with that, let's go!

Welcome to the podcast studio, virtually at least. Hi Jürgen, great to have you with us.

Jürgen:

Thank you, Madeleine, and thanks for the invitation. I'm glad to be here and to shed some light on this whole area of AIoT and industrial AI – at least on what we're seeing in the market, what we're seeing at customers, and how Cumulocity is positioning itself. There's plenty to talk about.

Madeleine:

Lovely. I'm looking forward to it too. I've been waiting for this episode for ages, and I think a lot of people out there have as well. So it's great that you're here. Quick one first: virtually – where am I reaching you right now? Are you in Germany or somewhere else?

Jürgen:

I'm in Germany, in the office as usual, all good.

Madeleine:

And where is your office again, for everyone listening locally?

Jürgen:

In Düsseldorf. We have several offices in Germany, but Düsseldorf is our headquarters.

Madeleine:

A shout-out to Düsseldorf then, and of course to everyone listening, wherever you are. Jürgen, a quick introduction for anyone who doesn't know you yet. As I said: CPO and Managing Director at Cumulocity. Since 1 January 2025 you've been an independent company again, after a management buyout from Software AG in 2024. Cumulocity operates in more than 20 countries and over 1,000 companies use your platform. I've known you for a long time now. And I believe you've connected millions of devices by this point. Maybe give us a quick sense of the company, so everyone knows where we're starting from.

Jürgen:

That's all correct, and we're proud of it. These are industrial devices – that's the point. There are IoT providers who connect more consumer devices, lawnmowers or streaming boxes for Netflix, and those come in much bigger volumes. But our focus is industrial IoT. When we talk about devices, we mean wind turbines, elevators, compressors, pumps. We have more than 25 million of them connected and running through the platform. Which also tells you we've been in this market a while and have built up a lot of experience.

Madeleine:

I was about to say – I think a lot of people already work with you. And for those who don't yet: stay tuned for what we're about to cover. I'm glad you're back on the show. We've done an episode together before, the last one was in 2025 I think, so it's good to get an update today. There's also an episode on the Cyber Resilience Act – if that's of interest, I'll link it in the show notes. You can hear there how Cumulocity approaches that topic. Today, though, it's about AI.

Maybe it makes sense to put the topic in context and define it a bit. I'll start from my own perspective and you can add the practical side. Because there are different types of AI. There's classic machine learning – people from data science, but plenty of others too, have known that for years: camera-based quality inspection, reject part detection, all of that. Classic machine learning has been around for decades.

What has arrived more recently is generative AI. I think many of us work with ChatGPT, Gemini, Claude and so on, where you can actually talk to the AI, have things explained, summarise texts, but also get real recommendations and interpretations. And then there's hybrid AI, which combines the two. That's been my high-level understanding of the AI landscape out there, let's say.

Jürgen:

That's exactly right. We make the same distinction between those two flavours of AI. First, classic machine learning, which is used more for analytics – you train machine learning models on machine data, for anomaly detection, for the remaining useful life of wear parts, or for object recognition in images. That's the classic machine learning use case.

And then, for some time now, there's agentic AI and the LLMs – the route where you interact with an AI through natural language and put that to use, for example to query data from an IoT platform or to build applications. We'll come to that, because these agents can do a great deal now. It's quite impressive what's possible.

And when we talk about AIoT, it's really where AI and IoT come together. It's the step from classic machine connectivity, which we've always done – connecting the OT world, lifting the data through the edge into the cloud – to connected, autonomous processes that use AI to push automation and autonomy a step further. That's what AIoT means.

So both types of AI are in there: classic AI, machine learning, as well as the newer approaches in generative and agentic AI. We can go into examples in a moment. But AIoT simply brings AI and IoT together, and IoT is the precondition, because you need the data first – otherwise there's not much to analyse on top.

Madeleine:

Exactly. Would you say AIoT always means live data from those devices and machines? Because I've been wondering: customers run very different projects. Sometimes the data sits in a larger data lake, sometimes you go straight through to the device itself. Is it always live data, or is it sometimes just historical data sitting in a database somewhere? Would you always talk about live data?

Jürgen:

Classic machine learning models are of course trained on historical data. You're quite right: many customers consolidate their data landscape in data lakes these days. And IoT data is an important source. But there's plenty of other data in a company – business data and so on – that ends up in those data lakes as well.

What we're seeing lately is that the data in a data lake or lakehouse is getting fresher and fresher. We have a new component in our DataHub 2.0 that gets data from the machine into the data lake within five minutes at most. And then you run your classic analyses on it. In the past that would have been a BI tool sitting on top, something like Power BI. Today it's machine learning models as well – someone has a Jupyter notebook or Python and works on the data. Or, very much of the moment, agents access the aggregated data and the context and try to optimise business processes. That's possible too.

And as I say, it's interesting to watch, because in the past this was strictly separated: everything real-time was operational, and then you had historical data in your data lake that was usually quite old, a few days or weeks. That's disappearing. We're down to five minutes as the upper limit now. In practice you're under five minutes – your data is in the lakehouse within one or two minutes. Which means you have more or less fresh data available for your AI agents and applications at all times.

Madeleine:

Fresh data is a good way to put it, because it isn't live data or real-time data – it's simply the data you work with. And what you do with it in the end always depends on the use case.

Jürgen:

Exactly, that's right. And you still have the edge dimension as well. We're very strong in streaming analytics and live processing. If you're taking labelled images from a thermal imaging camera and you have to decide quickly at the edge whether something is overheating or an anomaly has occurred, so you can raise a warning or an alarm – then of course you do that with streaming analytics rather than sending the data to the cloud first and seeing what happens.

So depending on the use case you have to ask: do I need to do this at the edge, do I really need streaming analytics because I want to be on it as fast as possible, or are a few minutes enough? For many applications, a few minutes are perfectly fine.

Madeleine:

Absolutely. You're already deep into the use cases. I'd also love to talk about the different agents later, and I know data harmonisation is a big part of this – a lot of that gets easier with AI, and I think we'll come to it today. But before we go there, let's pick a project out of your portfolio. You have so many different customers. If there is one already running in practice – just so we have something concrete underneath our discussion. Is there a project we're allowed to talk about?

Jürgen:

Yes, I gave that some thought beforehand and I've brought two projects along. They don't map perfectly onto agentic AI yet, but both have a flavour of it, so you can see where this is heading. Because a lot of companies – and we should be honest about this – are looking at the topic and are at prototype stage. Which is understandable.

One example I have is Vision AI, and that sits more on the classic AI side. We worked with SAP on it. SAP uses Cumulocity inside a business application called Asset Performance Management, APM. Cumulocity sits under the hood there, providing the connectivity to brownfield and greenfield machinery and bringing the data in. The data is then mapped onto the SAP asset model. That all happens out of the box, the customer never sees it. After that – as the name says, Performance Management, essentially predictive maintenance – various KPIs are calculated on that data and actions are triggered, for instance notifying a technician when there's a problem with a machine or maintenance is due.

And the use case I want to show today is visual inspection of industrial equipment using a thermal imaging camera. We work with Sony on this, for example, though it could be another manufacturer. The camera is connected to a gateway – that can be a Raspberry Pi – which in this case sends the data to the cloud via our thin-edge.io.

Madeleine:

That's one of your products, will we come to it?

Jürgen:

There's an AI model analysing those thermal images, and it runs directly at the edge. It's genuinely lightweight on thin-edge.io. What you want, of course, is to spot unusual overheating in real time. A thermal image looks like a colour gradient, from deep red through yellow and orange and so on, so you can picture what I mean. The AI model labels the image, looks at it and works out how hot the object being measured actually is. And when a critical temperature is exceeded, the system automatically sends the detection data through Cumulocity to SAP Asset Performance Management, so the maintenance team can act on it.

And what we do there is train the model – depending on the use case. It might be thermal images of a motor, a bearing, a control cabinet, it depends. Or in quality control there are joining and welding processes – laser welding, plastic welding, adhesive bonding – where you use Vision AI to track temperature profiles of the seam. Or with gas and liquid leaks, where infrared radiation is absorbed. A Vision AI approach makes sense there too.

And the training happens on the historical data, on the data in the lakehouse. You normally have long-term data there covering the last few quarters. That's where the model is trained.

Madeleine:

I was just about to ask where the training data comes from.

Jürgen:

Exactly, that's where it's trained. And then we take the trained model and push it right out to the device. Either onto the device itself or onto thin-edge.io, so close to the device, and we evaluate the image data there in streaming analytics. That's called inferencing. You take the measurement data, the labelling from thin-edge.io, it goes into streaming analytics, certain parameters get checked. And once the threshold is crossed, an event is generated that travels up into the platform.

Madeleine:

Two quick follow-ups. For anyone who has never heard of thin-edge.io, perhaps because they don't work with you yet – could you explain briefly what it is? And what exactly does "back onto the device" mean? Onto the thermal camera, onto an edge device – or is the camera itself the edge device? Two quick questions.

Jürgen:

So thin-edge.io is our open source framework for getting data from a gateway, a PLC and so on into the cloud. There's a broker in it, so you're fairly flexible downwards in terms of data connectivity, protocols and support. And then there's a messaging broker that essentially takes the data and pushes it up into Cumulocity. But it can push to other clouds as well. It's an open source framework, it isn't tailored to Cumulocity – it works perfectly with Cumulocity, but it can also put the data into Azure or AWS.

Madeleine:

So an IT/OT integration layer, essentially. It collects the data and passes it on.

Jürgen:

Exactly, a bridge from a gateway or a PLC into the cloud. That was the first question. And on the second one, what I mean by action: it depends on what's being used in the given case. If the camera is powerful enough to run the model itself… And with a Raspberry Pi there's a camera plugin where you simply plug the camera in and the whole thing – the model and the inferencing – runs on that Raspberry Pi. So it varies. Sometimes it's a retrofit and it runs on the gateway instead. You have to look closely at how it's set up. But the point I wanted to make is that the footprint needed is small. We're talking a few gigabytes of RAM, no more. It doesn't take much.

And then you do what's technically called ML Ops. New data keeps coming in from my machines, it goes into my data lake, and I retrain the model. So there's a continuous improvement loop: new data, retrain, get a better version of the model, and deploy that back down to the edge. Cumulocity supports that out of the box as well. We handle it through our device management functions – a machine learning model is simply software with a newer version number. That's how it's implemented.

Madeleine:

Nice. And in this use case, the AI stays at the recommendation level for the people who need to know. It doesn't intervene autonomously, does it? That would be the next thing, which is really what we're here to talk about – AIoT. Sorry, I may have jumped ahead.

Jürgen:

That's the second step, exactly. And that's the interesting part. At that moment the classic machine learning generates… I normally have high-frequency data here, new readings every second or several times a second. I don't want to run an LLM on that, my token costs would explode. So I use classic machine learning for that. And then we generate these events – an anomaly, for example: the temperature is too high, the thing is glowing red. That's the event. So I know the camera detected that event at that point in time.

And what I do then is feed that into an AI agent. Now an LLM comes in and performs a root cause analysis, because it knows the camera sits on this motor, in this piece of equipment, in this factory. And there may be further context: it's a pump or a motor from this or that manufacturer, there's a service contract or there isn't, and so on. The LLM factors all of that context into its interpretation.

And then the LLM could – and we have prototypes for this, though it isn't in production yet – give the user a pointer and say: you probably have a problem here, this looks like the cause, here's my suggested fix. And then the specialist can say: understood, let's do that. Or: no, you're completely wrong, we'll do it differently. A human still makes the call.

Madeleine:

So we have the classic ML approach, then the large language model as the next stage. And AIoT would be one step beyond that – a device or a machine deciding for itself. But that's the future, probably five years out, isn't it?

Jürgen:

Yes. Technically it's possible, I have to say that. Technically it's possible. We're currently building fairly strong auditing and approval mechanisms into Cumulocity, because it's critical to be able to verify clearly whether an AI agent made a recommendation and whether the user reviewed and approved it.

Looking ahead, when people talk about the industrial AI trend you often hear terms like zero touch industries or dark factory. You'll have come across those – everything running autonomously. And that is where people want to get to with autonomous operations. But that's a way off yet. Before we fly a factory blind, it'll take a while.

That's why this approach – human in the loop, as it's called technically – is the right one for now. The AI agent is used to do a lot of the legwork. It goes through the log files, prepares everything for the person who actually has the expertise, and they can look at it and check it. But the human still decides whether something gets stopped or replaced.

Madeleine:

That matters. I think the way you describe it, the alarm bells are already ringing for a lot of people, because security is a huge issue here. I know from our community that people are restructuring their security setup right now. There's a lot of build-out work happening before anyone touches this. And you do have to separate hype from what's genuinely ahead of us. But it's good to see that projects like this exist and how they work.

Jürgen:

I did bring a second example, though. That one involves genuinely autonomous actions.

Madeleine:

One more question before I forget. You mentioned tokens. For anyone who isn't deep into this – can you explain when I need a token and when it starts costing me money? I imagine plenty of people who've launched their first pilots need to understand this, and some are building a business case. Can you explain where AI really starts costing money and where I need to be careful?

Jürgen:

Sure. Many people will know it from their private use as well – you've probably got a Gemini or an Anthropic subscription, and there are different plans depending on what you do with it.

And in Cumulocity it's set up the way it is because we think this matters, and our customers tell us the same: they want to keep control of their data and of their AI models. There's a lot of IP in there, intellectual property. What you want to avoid is some vendor training its models on the operating data of your machine fleet or your plant, building up IP from that domain knowledge and then perhaps selling it on to other customers. That's the risk.

We play this very cleanly. We're one of the largest European vendors, and data security, governance and trust matter a great deal to us. We understand the concern completely, and it's where we want to differentiate ourselves. That's why you can configure the LLM in Cumulocity. If you go into Cumulocity today there's a component called Agent Manager. That's where I set up and configure my AI agents – which data they're allowed to access, and so on. And I set my LLM there too. So a company can say: we're standardising on Anthropic or Gemini or whatever across the business. IT keeps control of that and manages it. And then that LLM is used for the agents in Cumulocity as well, with Cumulocity as the platform for those agents.

IT always likes that, because everything stays inside their predefined, clean framework. No shadow IT, nobody spinning something up on their own with data leaking out. And that's very, very important.

And in the example with the heat pump, those events: as soon as that LLM starts researching the root cause – why did this heat build-up occur? – the AI model goes off and queries APIs. Past data from the machine, perhaps the context from our Digital Twin Manager, where it can see that the pump or motor sits in this line, in this piece of equipment, in this plant, and was last serviced on this date. Those are all API calls and MCP calls – Model Context Protocol is the other term – through which the agents pull that context. And every time they do that and evaluate it, they burn tokens. That's the metric the LLMs use. And it costs money.

Madeleine:

So if I want to roll out an AI capability internally across our plants, I need to think through what that architecture looks like. Or how I handle this… the Agent Manager you already have – I don't know how many companies are currently building their own. That's exactly your case, to say: have a look at this. But I'd need to be extremely careful about how I set it up, because it can get expensive fast if I don't think it through, right?

Jürgen:

Exactly. And above all you have to think it through from a security angle: which agents have access to which data, and where are the models actually trained? Do they run off into some American hyperscaler I can't control, or do they stay with me? Those are the questions to ask.

Madeleine:

And also: which data do I hand to a manufacturer, where it makes sense to? You mentioned the thermal camera example – do I share that with the camera manufacturer or not? That's a question you have to work through: when does it make sense to share, perhaps for larger installations?

Jürgen:

Data sharing too, yes. And the agents… applications are moving more and more towards agents. It's worth seeing this as a transformation: in the past, applications were SaaS applications, lots of user interfaces where specialists clicked around, defined their rules, and had their dashboards. All of that is changing because of AI, so in future I won't just have human users on my UIs but increasingly – probably predominantly – agents querying my APIs and MCPs.

I don't think the cockpit or the UIs will disappear entirely. You'll still need some of it, I think, to visualise and monitor the most important operational processes with KPIs. But the shift towards agents as the main users of an AIoT platform will be very significant.

Madeleine:

Interesting. By the way, if you're listening and thinking: we're working on a project like this, or we want to – Jürgen, I'll put your contact details in the show notes. Have a look at what Cumulocity has built here.

And a quick invitation from our side: we run a user group. If you'd like to compare notes with others first and see how they're handling it, just come along. We meet monthly, users only. Drop in and join us. I'll put it all in the show notes, or have a look at iotusecase.com. I'll include the Cumulocity page as well. It's worth a look, and it's a good shortcut – you may not need to build this yourself. It's also worth asking yourself what you have today and whether it scales for AI use cases like these. Sorry Jürgen, do add to that.

Jürgen:

You could also include the link to the SAP Community. This Vision AI example is documented there, and it goes deep technically – for people who want to see exactly where the data flows and how the models are deployed. I can provide a link where you can read it up in full detail.

Madeleine:

Good point. Okay, before we get to the other agents and use cases – because we still want to cover how data harmonisation works. There are plenty of people putting a lot of manual effort into harmonising data. We'll come to that shortly. But you had a second example, and I didn't want to cut you off.

Jürgen:

Yes, this one is about autonomy and autonomous operations. We just talked about being cautious with giving agents full control and letting everything run autonomously. But we have customers who have fully automated parts of it, even if those parts are small. And Enercon is a great example.

Some quick background: Enercon develops, manufactures and sells onshore wind turbines – you'll see plenty of them across Germany. But they operate worldwide as well, with turbines in Alaska and even at the South Pole, in Antarctica. More than 30,000 wind turbines globally. They're a major customer of ours and the volume of data is staggering: those 30,000 turbines send in 500,000 data points per second. 500,000 data points per second – that's 43 billion data points a day coming into the platform.

Madeleine:

My goodness, and from equipment like that. There's no real infrastructure out there, they're standing in the middle of nowhere.

Jürgen:

Exactly. And all of that flows into a lakehouse, but through Cumulocity as well. If you look at the numbers: one day of downtime on a turbine costs around 8,000 euros on average. Enercon has a remote service team – several hundred people working in shifts, making sure 24/7 that the turbines are running and that when something happens, they try to resolve it remotely first. If they can't, someone has to go on site, and that gets expensive.

To give you a sense of the scale: they get several hundred thousand incoming service calls a year and several million SCADA alarms. It's a huge platform. And automating workflows is essential there, simply to save money.

What they did was enable their remote service users – the people sitting at the screen monitoring operations – to simplify daily routine tasks using our Analytics Builder. That's a self-service tool where they can model analytical logic without writing code. And we had an interesting use case where they restart a turbine based on a clearly specified pattern. Until then, a ticket would come in and a person would genuinely check whether certain parameters were met, whether everything fitted. If it did, they pressed a button and the turbine restarted. It was a deterministic pattern.

And we've now modelled that in Analytics Builder. The machine checks whether the parameters and conditions are met, and if they are, it triggers a reset of the wind turbine through Cumulocity. That's an automated remote action. It isn't AI in this case. There's no probability involved, no room for interpretation – it's a classic check against a set of parameters, and when all conditions are met, it acts. But it's still a good example: it let them automate around 40 percent of their turbine restarts, and the savings run into the millions.

It shows the power in that… which is why I brought the example. It isn't applied AI as such, but you can see how much power there is in automation. That's why I think you have to move there step by step. And as you get closer, a lot of value can be created. So I do think this direction – industrial AI, connected and automated operations – is the right one. You just have to make sure everything is traceable and works cleanly.

Madeleine:

Absolutely. As a company you first have to take the decision that the person who has been doing this until now… that you're automating it. That's already a bold step. But when you see the business case it makes complete sense, because that person can turn to other work – this was a recurring problem after all.

And it's interesting. I imagine a lot of people out there first have to find the use case where they dare to go down this road. You presumably don't do it in every use case, you have to be very careful. But there are cases where it makes sense. I find that fascinating.

Jürgen:

Exactly, that's how it is. You start with peripheral use cases where, if something does go wrong, it isn't catastrophic straight away. Which is only logical.

Madeleine:

Do you see a concentration of these use cases in particular industry segments? You go to market vertically as well, you have customers across all kinds of industries. Are there use cases based on or coupled with AI that already work, where you'd say there's a focus? Is something crystallising, or is it still too early to say?

Jürgen:

I think in distributed energy resources – energy and utilities – there's a lot going on already. The pace of innovation there is quite high.

In manufacturing there are early use cases too, and we're working with a number of customers. But our focus isn't so much on the shop floor. We're stronger when it comes to connecting large distributed fleets than in a single shop floor automation use case. Beyond that, we have quite a bit in medical technology and defence at the moment, although defence is of course very conservative. There it comes down more to being able to deploy in a sovereign cloud.

In medical technology the focus is currently less on agentic AI and still more on OT/IT connectivity, security, certifications and making the data accessible in the first place. That's the crucial step. It's great to have AI agents on top collaborating and automating further. But the crux is this: without a securely connected fleet there's no point starting at all. That's step one. If the fleet isn't securely connected, you may as well stop. After that it's about data preparation. If I feed bad data into my AI, I get garbage out.

Madeleine:

Exactly. I'd like to close with… you're already starting to give real advice and share what you know. Maybe we can pull that together at the end – what your closing advice would be for companies working on this.

But one question before that, to summarise what we've covered. You've named all sorts of use cases on the AI side. Can we pull together the top five use cases where you'd say: these are the ones that already work? Just so companies can assess for themselves whether they could do something here, or have already done it. Can we summarise them?

Jürgen:

From a Cumulocity perspective I'd lean towards horizontal use cases. We do have Vision AI, and in energy and utilities there are genuinely domain-specific use cases. But we run an AIoT platform, so we asked ourselves: what can we put into our customers' hands?

And one really interesting use case is data harmonisation – you touched on it earlier. It's an age-old IoT problem. IoT is extremely heterogeneous: all sorts of protocols, old devices, very old devices, brand new ones. You simply have to live with that zoo of devices and protocols and make the data usable. And we've now turned AI loose on it.

What we have is technically called an Expert Agent, an AI-driven component in Cumulocity. What it does is listen in on a moment of live data on an incoming data stream, a messaging channel, look at the raw data, and then propose to the user how to transform it into the Cumulocity format.

In the past that always meant coding – unless you were dealing with a newer protocol like OPC UA or MQTT. And because the payload often differs, a person always had to get hands-on before the data landed cleanly in Cumulocity. Every platform has this, it isn't specific to Cumulocity: all kinds of data arrives and I first have to make it usable. And an agent does that now.

Madeleine:

Just so I've got it: in the example you gave earlier, the thermal camera – you'd have one camera, or several networked across a hall, and then others in a different plant. And that naming, that mapping, was largely done manually: this is the data point, it's called this, it's probably a temperature sensor. That naming of data points is what used to be manual?

Jürgen:

Exactly. The naming of data points is one part, the units are another. And how is it transmitted at all, if I have an image? How do I represent that image? Sometimes as a vector, so I can write values into different columns. And every manufacturer does it differently. You have to cope with that.

We've tested this in practice with a range of formats: OPC UA and Modbus of course, but also LoRaWAN, Protobuf, CBOR, Sparkplug, JSON and so on. And the good part is it works remarkably well. In the product it proposes a smart function, a piece of JavaScript, for transforming the incoming data into the Cumulocity format, and it also lets you debug and test it. That saves a remarkable amount of development work. It's a case where we've taken an old IoT problem that always needed hands-on work and revolutionised and automated it with an AI agent.

Madeleine:

Interesting. Does that apply to rolling out edge devices as well? If I have several edge devices – I've actually got a request in our network right now – you have to apply different configurations. Is there a way to push configurations automatically onto a new edge device? Just a different angle.

Jürgen:

Yes, we're working on it. It's a hot topic in our R&D right now, I can tell you that, Madeleine. A software rollout agent is on our roadmap, it isn't in preview yet. Cumulocity can already do part of this: you can define device profiles and we support bulk updates. But it's always manual, someone has to click, and then a process runs in the background that tracks whether it succeeded or needs another attempt.

What we want is to automate that with an agent, so you can simply say: here's a new software version, deploy it to all devices of this type in Europe or in China. And it does that in the background and reports back whether it worked or where there are problems. We're on it, definitely.

Madeleine:

Okay. So that's what you meant: data harmonisation first, which is about mapping the data from a sensor, say. Or into an asset management system – the onboarding, essentially. Would you put it that way?

Jürgen:

Exactly. What kind of measurement point is this: a temperature sensor or a pressure sensor, is the unit Celsius or Fahrenheit? In one plant the sensor is called Hugo, in the next it's called Willi. Naming out there isn't standardised. You can set it very freely on the PLC, and unfortunately people make full use of that.

Madeleine:

Exactly. It's remarkable. I've done so many episodes on the Asset Administration Shell and so on, all the approaches out there.

Jürgen:

Yes, or the whole Unified Namespace topic, UNS for short. Exactly, that covers the same ground.

Madeleine:

A fascinating subject. So if that interests you, just follow the podcast – there's plenty on it.

Jürgen:

The other good use case is AI-assisted development. In Cumulocity today… we used to have a cockpit where you clicked your dashboards together: I'd like a line chart here, a status indicator there, and a speedometer over there. You can do all of that through the agent now. You just tell it: build me a dashboard for this use case, put this and that on it. And it knows which data points to pull from Cumulocity, how to wire them up, and it builds the dashboard for you. No more coding or clicking, the agent does it. Those are typical use cases that work everywhere, across all domains.

Madeleine:

That's a good one, and a typical use case. So we have harmonisation as one use case, with agents for it. Then AI-assisted development. Do you have three more? We've touched on others already. Maybe we can get to three.

Jürgen:

Well, the rollout agent is in progress. We're also working on IT/OT security. The downside of AI is that it finds vulnerabilities faster. You'll have seen what's been happening – agents have genuinely found security holes far faster than people would have liked. The upside is that you can use agents to close those holes quickly too. So you can picture an Expert Agent for IT/OT security that keeps the fleet up to date, makes sure vulnerabilities are tested, and so on. There are all kinds of ideas we have for Expert Agents in Cumulocity.

Another interesting and important area is assembling Cumulocity data for external agents. Cumulocity is a great platform, but we know that for these business processes we're only one player among many, one of many systems that need to be connected. So there are external agents that need access to the data. And we're investing heavily in providing a strong MCP server, so that a Joule agent or another external agent on the customer side can pull data from Cumulocity together with the context.

Because the context is what matters. Raw data on its own is just noise, it gets you nowhere. You have to know: this is a temperature sensor, it measures in Celsius, it sits on this machine, in this hall. Without the context you're lost. That's why providing that context is so important, and in Cumulocity it happens through the Digital Twin Manager and then through the MCP server.

Madeleine:

Okay, good. And you have a whole network on your side, since you've mentioned external data sources and connection points. You have a large ecosystem there. So you still work with partners as well. When it gets into specific use cases, I can connect what I need. You mentioned SAP as a partner – there are probably many more.

Jürgen:

Yes, Microsoft on the OT/IT security side, and SAP, a very strategic partner for us, where we put the data into the Business Data Cloud and SAP Business AI with Joule sits on top, so we can feed the SAP agents as well.

And the decisive part here is that raw machine data without context is useless. You always have to close the gap between OT data and business context. That works rather nicely in the collaboration with SAP, because we link the SAP asset models directly to the Cumulocity Digital Twin Manager. The OT world – reality, let's call it that – always looks different from a logical model in the ERP. What we do is link the two in the platform, so the agents see the logical world, how the asset is represented, while we handle the complexity underneath: data coming from the historian, data coming in via a LoRa sensor and so on. Cumulocity takes care of that.

And with that context, the agent can genuinely do something. It understands what it's supposed to do. Otherwise it just fires wild queries at the APIs, starts hallucinating and burns a lot of tokens. And that's a lesson worth taking away: this context layer, the semantic layer as we also call it in English – that's where the real value sits. And it's domain-specific. Which is why we need domain-specific partners, because the semantic layer looks different in manufacturing than in energy and utilities. That's only natural.

Madeleine:

Do you have an example of what that context layer looks like in practice? Just so people can picture it. Maybe with the thermal camera – does that work as an example?

Jürgen:

Yes. What I mean by semantic layer is the metadata of the data points. This is a temperature sensor in Celsius. This is a pressure sensor. This is a thermal camera delivering images in this format. That's one part. Then you have the asset hierarchies and relationships. You know exactly where the sensor is installed. Because how else would the AI agent know? It sees some measurement values, which look odd at first, and they change over time. Without context it's more or less helpless.

But if it knows the sensor sits in this component, in this machine, in this plant – and the other sensor sits right next to it. Or take the wind turbine example: two turbines in the same wind farm. That's context. Or the product documentation with all the fault codes: if a pump throws error 4711, what does that mean? If I at least know the pump is throwing an error and here is the documentation – the agent reads the manual. The technician might not be keen to, but the agent reads it and finds the answer. That's what I mean by context layer. You have to model it and make it accessible.

Madeleine:

Of course. And you have a fantastic partnership with SAP. But as you say, there are all sorts of domains, so you need all sorts of partners to model that and build the context layer. Do you also do that with customers yourselves, going in as advisors and saying: we'll help you build this context? Or do system integrators and partners handle it?

Jürgen:

Absolutely. In some industries we know our way around very well – in DER, distributed energy resources, we've been working successfully with customers for years. Or in medical technology, where we've built deep experience. We do that ourselves. But even there, there are projects where we're on site with our own professional services and a partner alongside. In other areas we rely primarily on the partner, because they have the industry experience.

When it comes to autonomous business processes or optimising them, you need a partner who can do that. And Cumulocity would – to be blunt about it: if we walk into a Daimler plant in Sindelfingen and say we'll improve your OEE by 3 percent, nobody believes us. When we talk about connectivity at the base layer, people might still believe us. But walking in and saying: listen, we'll push your OEE up by 3 or 5 percent – we aren't credible there. With a partner we are. We're a technology supplier. The technology has moved on, we now have a bigger toolbox with IoT tools in it and a few AI tools as well.

Madeleine:

Strong. So maybe a final question for today – I teased it earlier, the advice. A lot of people are working on these projects and wondering what to do. Can you share what you've learned the hard way, where you'd say: watch out here, this is where you can really burn time and money? You've touched on some of it. Or where you simply have to be extremely careful. Could you talk us through that for the people out there?

Jürgen:

Sure, a couple of things. First, this context layer is critically important, because it's what lets the agents understand the key relationships and deliver sensible results. If you don't use it, you burn a lot of tokens and risk hallucinations. You can spare yourself that. You really have to think about how the agent can make sense of things – and that sits in the semantic layer. That's number one.

The second is AI-assisted development, or, colloquially, vibe coding. It's happening a lot right now, and you have to make sure you don't end up with shadow IT. You can make progress very fast, but in our environment – industrial IoT – you have to be conscious of what you're dealing with: business-critical processes that have to run 24/7.

I can stand up a prototype quickly with vibe coding or AI-assisted development. But then some people might start thinking: do I even need a platform? Can't I just write it all quickly, say in Claude Code: build me the solution. So I've vibe-coded the whole thing. That might take a few million lines of code. The problem is I have to operate it. And that's where it gets ugly. That's where you have to be extremely careful. You can't rely entirely on a vibe coding approach for smooth operations.

That's why we distinguish three layers. There's the system of record – that has to run stably, be secure, have governance and auditing built in. That's where we see Cumulocity. Then there's the semantic layer. We see Cumulocity there too in terms of tooling, but also the partners and the domain expertise, the people who know exactly how things fit together and what the business rules are. And on top sits the agentic layer – the agents that run analyses and optimise operations.

At the application level on top, AI-assisted development is excellent in my view. You can cut development time and move very fast. But down at the system of record you have to be very careful, because that's where you can break a lot.

Madeleine:

I can believe it. That's good advice. For everyone listening: have a look at this in your own operations. If you have questions, get in touch with Jürgen, come along to our user group where exactly these topics get discussed. And I'll summarise the top five AIoT use cases we've covered in the show notes. Have a look there, we have plenty of material you can read up on.

And from my side: Jürgen, thank you for your time today. I found it fascinating. We had two concrete projects, we defined things cleanly – classic AI, generative AI, hybrid AI and how it all works with agents. And what matters at the end of it. So all I can say is thank you, and I'll give you the last word. Great to have had you on. And perhaps we'll do an update in a year or so – we're in touch anyway.

Jürgen:

Gladly, any time. Thank you, Madeleine, for the invitation and for having me on the podcast today. Anyone who wants to try this out can go to our website, there's a free trial. You'll have it within a minute – you get a tenant and can start straight away.

And for anyone who wants to look into it and get a better sense of where the starting points for industrial AI might be in their own company: I'll offer everyone who has listened to this podcast a free day of consulting from Cumulocity. So come and see us, I'll show you some demos and we'll brainstorm together. And perhaps something comes of it.

Madeleine:

That's a generous offer. Thank you, Jürgen, and have a good rest of the day and a good rest of the week. Take care, bye!

Jürgen:

All the best. Thank you, Madeleine.

Have a concrete IoT project in mind?

We know the vendors who've already done it.