ESC
Proving resilience for insurance risk platforms
September 16, 2026

About this webinar
When disruption hits, can you still stand behind your risk numbers? This webinar explores how insurance organizations can move beyond IT uptime metrics to build defensible, evidence-based resilience across actuarial workflows, data pipelines and governance frameworks.
Speakers
- Neil Covington, Senior Principal, Product Management, General Insurance, FIS
- Terry Buchner, Principal Insurance Specialist, Amazon Web Services
Duration
29 minutes
What you'll learn
- Why resilience in actuarial risk platforms is a credibility issue, not an IT uptime metric, and how silent data errors can be more damaging than system downtime.
- How to apply a continuity, recovery and control framework to protect critical actuarial workflows, such as reserving signoffs and capital modeling under disruption.
- What auditors, regulators, boards and shareholders expect as proof of operational resilience and how to build evidence-based artifacts that satisfy each audience.
- A five-step practical path to mapping actuarial workflows, defining impact tolerances and embedding resilience as a continuous business discipline.
Expand to view full transcript:
Neil Covington: Welcome to today's Tech Talks, Proving Resilience for Insurance Risk Platforms. Thanks for joining. Today is not yet another AI talk. or a cloud architecture session. It's a risk governance conversation, told through technology. The question we want to explore is simple. Can your risk decisions withstand disruption? We're going to talk about resilience, the way actress experience it, as defensibility of decisions under pressure. I'm joined today by Terry Buechner from AWS, an industry expert when it comes to cloud platforms. Terry, let me start with the hook that's been resonating. It's not, is the platform up? It's can we still stand behind the number when conditions degrade?
Why is that the right framing right now?
Terry Buechner: Thanks, Neil, and first of all, thank you very much for having me. Very excited for the talk that we're going to have around this. It's very cool the way that it's been framed, I think. I think maybe interesting or surprising to have somebody from AWS and not talking specifically about IT or AI or other things.
And that's very important, because we're talking about it from all this resiliency risk from a business perspective, which is, as you said, how our teams actually experience it, right? And I think the cool thing is that, or interesting, is that resilience for those teams, for actuarial teams, risk teams, it's really a… you can think about it as a credibility issue. It's not an IT uptime metric. And when I say credibility, I'm talking about the numbers at the end, these key, important numbers, capital, reserves, stress scenario outputs, things like that.
And if you think about it, why it's important not to think about it purely in the sense of IT, although that is an important piece of it, of course, it's because how we came to these numbers at the end, it's from a long data supply chain, and it is truly supply chain. It's not a physical product in this case, but it is a supply chain, and this supply chain has been built by our carrier customers over, in some cases, decades, right? If you think about all the policy administration engines, administration systems, excuse me, that are in there, which are… can be manifold, at least more than a few for more, most carriers.
Claim systems, external data, and it's super important to be ingesting the vast amounts of available and more of data becoming available every day. If we're not taking advantage of that data and not ingesting it in a correct way, then we're not getting a full picture of risk overall. But when you get that data, whether it's internal or external, you need to enrich it, think about ETL pipelines, all of that.
And as, as I said, we built… we had carriers have built these chains, these, supply chains, incrementally over decades for them, right? So that… and it's through acquisition, through buying, through building. So when disruptions happen, and they do happen, whether it's a major one or a minor one. Often, although the disruption itself is certainly a big issue, and there's problems around that, but it's often the subtly wrong results, right, things that you don't know are wrong, or slightly off, and they're not obvious, that can be caused by partial or stale inputs that cannot be seen until almost it's too late, in many cases.
Neil Covington: Yeah, I think that's the key point, Terry. Wrong and unnoticed is worse than down and obvious. And we're seeing this becoming more than best practice, more like expectation.
Terry Buechner: Yeah, I think from the expectation there, and in insurance, it's interesting, I think that because we're… although there's a lot of interesting technology, and we've been using technology for a long time, almost since the beginning of the insurance industry. it is… we're a bit slow moving in these things, and a lot of the pressure comes from outside of the insurance industry. So, when we think about how regulators, auditors, shareholders, boards, policyholders are expecting us to not just say that Something is true, but rather demonstrate these things, often in real time, or more in real time as well.
So demonstrating operational resilience is very important, and you think about what regulators are expecting us to be able to provide to that, the bar is getting set higher and higher all the time. The other side of that is that the tools are available for us to do so. And part of that is modernizing the data state, right? So we see many insurers modernizing their core systems, modernizing their data states. Moving from a fragmented data state to a more unified platform, and that's for a variety of reasons. It can be business-driven, meaning growth, or profitable growth in terms of modernization, business agility. So, all those things tied together, and tied together very nicely if we think of them outside of individual silos. So, it's a modernization move, yes.
It's an IT resilience move, yes, but it is a overall risk resilience move. If you think about having fewer point-to-point integrations. a clearer data lineage. So when something breaks, well, first of all, you have fewer vulnerable points throughout, if you have a unified approach and uniform workflows. But you also have a clear view of what broke, when it broke, and be able to go back into a point in time and try to fix that.
Neil Covington: It sounds like the transition is, if resilience is a risk control, not a checkbox, what does that actually mean in actuarial terms? So if we break that down, using the triad we've been using, continuity, recovery, and control, perhaps, first continuity. What should actuaro leaders be doing differently, do we think?
Terry Buechner: I like this triad approach, having sort of a very clear, mechanisms around that. We're very big at mechanisms at Amazon, so I think having mechanisms around that and definitions is very important. So if you think about continuity. I like to think about it, and I've seen carriers successfully think about it in terms of decision windows. So, if you think about… and not the systems, although, again, important to remember that we're not saying the systems aren't important, they are very important.
But thinking about this from the business perspective. So, from continuity perspective, we think, what workflows, which ones of my workflows must still run, no matter what? Maybe it's reserving sign-offs, or calculations. you know, there are deadlines, hard deadlines in many cases associated with these ones. So, identifying those. And thinking about what those windows are is key to that. Some other analytics could be business intelligence or other things, maybe they can pause. So defining the criticality of them and the defining around that is super important around this, and not which server is the most important one to bring up, because it has our core policy system on it. Rather, what is the… what is the decision deadline that we have to sign off on something and think about it in those terms?
Neil Covington: So it's more workflow-based, not infrastructure-based, would you say?
Terry Buechner: Yeah, absolutely. So, it might be funny to hear from a cloud infrastructure provider to say that. The infrastructure is very important, of course, because it's the systems that support all these things. But we need to work backwards from what we're trying to achieve from a business perspective. The workflows are the key here.
And I think the data integrity is that, as we talked about a little bit, the hidden resiliency risk, right? So the data integrity, if a feed is delayed, or maybe it's partial, maybe the data is stale. So, perhaps the platform can still run, because there wasn't an obvious error in it. So the platform… so the workflow is still running. But the output maybe is misleading at that point. So the risk isn't the downtime itself, although that is a risk, but it's also that, what we would call, and we talked about this as the idea of these silent error propagation, something you don't know that's happening and can continue and have knock-on effects downstream.
Neil Covington: Okay, so if we think about a concrete example, perhaps, so say it's the morning of reserving sign-off, and a key claims feed is only partially in. What does resilient look like?
Terry Buechner: I like to think about it in terms of having a predefined, degraded, but controlled mode, and you can call that different things, but the data is degraded, the workflow is degraded, but we're still in control of it, we understand what's happening. So, do we decide, again, based on those, on those decision windows, do we decide that we run on the last good snapshot? Do we apply a conservative adjustment to it, or do we have to pause and wait until it's fully back up? Very important to that is who has the authority to decide, that. We need to identify that, have it very clear, have backups to that, too. Perhaps somebody's on vacation. We need to be able to capture that, capture those signs off… sign-offs, make sure it's very transparent, so we need to be able to defend to auditors, to the board, to regulators, why a decision was made at a certain point of time.
Maybe it's, do you know which reserve segments are affected by the missing data? Can you produce a result for the segments that you do trust? Flag the affected segments, etc. There's a lot of pieces around that. Get conditional sign-off, maybe. Degraded but controlled. So we understand what's going on, have very… we're clear about the decisions we made around that, we have control over it, not just run it and hope for the best, and hope for that things will work out at the end.
Neil Covington: And people will typically hear the word recovery, and what they actually hear is RTO and RPO, and immediately think IT.
Terry Buechner: And that's understandable, and I think the interesting thing is that in the… if you go back in time, sure, it began with a business process, thinking about RTOPO from a business perspective, but the underlying IT supports that, and over time, in my opinion anyway, is that it's become more of an IT issue rather than… we've kind of lost the original business decisions that led to that.
So we need to think about it in business terms. We're going to resurface that, really think about that in business terms, and understand the IT that supports that as well. Again, those decision windows, the reporting cutoffs, any tolerance for delay that there might be, whether that's internal, from a risk perspective, external to our customers, or the regulators around that, too. So it's not… restore the database in X hours, but rather, what's the latest point that we can resume, but still make that critical decision on time? Or how much can we delay that and tolerate before the output becomes misleading? Or can we pause this because it's not a critical workflow that needs to be brought forward until we're fully back up and running?
Neil Covington: And then the third leg, the control, that's one that's easy to overlook, I see.
Terry Buechner: easy to overlook in a critical situation. It happens to everybody in all sorts of situations, but when you're in that disruptive period, which of course always happens at, like, 2, 3 in the morning, you're under pressure. So there… there could be… and that pressure could be internal, external, a variety of things, but get that run out, we need to get sign-off.
But the control around it, access, change control, the data checks, approvals, audit trail, etc, all that is essential. And almost more importantly during these times of crisis, they can't be optional, right? Again, you can't just get the run out, you can't hope for the best. you know, you need them the most at these points. So that resilience needs that you have, the degraded but controlled situation going on.
And be able to control that and understand what's going on. And Neil, I think you know this by now, but I'm a volunteer firefighter, been doing that for about 12 years now in my hometown, and there's this concept that is applicable here, it's called… we call it freelancing, which means that you go off and do your own thing on the fire ground, in an emergency situation. Because either it seems like a good idea. Or, more often, and more dangerously, because you're in that crisis mode, something's happened. Something's happened, and you make a snap decision, and react on instinct. and hope for the best, rather than having that control of training and chain of command. And on the fire ground, that can, you know, in a worst-case scenario, it can get people hurt, right? But in best case, it's going to be self-optimal as it comes out.
In an actuarial workflow, those ad hoc workarounds, which are used very, very often during disruptions, they bypass the controls that give you, and give our… give the auditors, good regulators, etc, confidence in the output. And you may not discover that damage until it's too late, and you may not see that damage until way down the road, like during an audit.
So it might not be physically dangerous, like, on the fire ground, but it still can be dangerous to the enterprise, for sure.
Neil Covington: And kudos for being a volunteer firefighter, Terry, for sure. I mean, I wonder how many of the audience have an explicitly documented degraded mode for reserving or capital flow, capital workflows and so on, including who signs off the exceptions.
Terry Buechner: Very good question, and I think that needs to be looked at on a regular basis, you know, having the resiliency more than just once a year, thinking about DR, all those kinds of things, and I do think that That modern data platform approach can really help Making it more explicit, whether you're dealing with that missing data means that we stop, we pause, versus stale data means we can proceed, but we have some reservations or caveats around that, too. It's a control. It's a control that we have, and we have documented it's not an opinion formed at 2 in the morning.
Neil Covington: Absolutely. And a perfect transition, so… Having controls is necessary, but boards, auditors, regulators will ask the same follow-up, prove it. And from my experience, the… I've always been told, in terms of audit, if you can't prove it, it hasn't happened. Learning point two, it's perhaps the difference between believing your resilience and proving it, then.
So, we've been using a blunt line. If your only evidence is a DR slide deck, you have an assumption, not a control. Why does that resonate, do you think, Terry?
Terry Buechner: I think that we need to have that in place because different audiences are going to expect different, proofs, right? They're going to… the resilience expectations are shifting toward, excuse me, toward evidence, and the different audience want different types of evidence behind that. Kind of different motivations, if you will. And there are many different stakeholders, but some of the key ones I think about are, of course, the board. They want impact tolerances. What outcomes are at risk? What are the tolerance thresholds for the enterprise as a whole? Auditors, they want to see control design, they want to see operating effici- effectiveness.
Our regulators, of course, they want tested ability to stay within whatever the tolerances they've defined, and there are many different definitions of that, many different regulatory bodies and different approaches to that, too. Very important to remember, we have shareholders in many cases, not all cases, of course, with mutual, etc, but shareholders, they want that steady ship. They don't want to see chaos, they don't want to see that ship tossed into chaos, into the stormcast during a disruption. Or, perhaps even worse, being pinged by regulators because of that chaos.
And then policyholders, which can be in a mutual, the shareholders, too, but the policyholders, really, they just, for the most part, they want assurance that their claims are going to be covered, that no matter what, their claims are going to be covered in a timely fashion. So they… and they want evidence of this, too. That evidence for policyholders might be lack of being pinged by a regulator, but still, that is an important thing to that. So having that evidence is important and needs to be defined by the audience as well.
Neil Covington: So how do we convert resilience into things you can actually show?
Terry Buechner: Think that, you can think about it in terms of, Proof artifacts, if you will. Sorry, there was a… I thought I had turned off my announcements, but that came up, so let me just start back from, neil, can you ask that question again?
Neil Covington: Okay. I'll just pause for a moment. So, 3, 2, 1… So how do we convert resilience into things you can actually show?
Terry Buechner: Approach can be to have three proof artifacts, if you will. One is having documented tolerances, documentation is very important, and those key recovery objectives, but again, not from an IT perspective, but from a key workflow perspective. One is having the test evidence behind it, as we talked about a bit, scenario-based And one is having the ongoing monitoring and attestation. Make sure that the posture hasn't drifted, but be able to test and visualize these things on a regular basis, too.
Neil Covington: So let's take the first one, so tolerances. What does good sound like?
Terry Buechner: Great, let's try that again. Another announcement came up that I tried to turn off, I apologize, please. But I've turned them all off now, so please ask that question. Let me just do one more thing, sorry. It's off. Should be good. I apologize, please continue.
Neil Covington: That's alright. You could hurry up. Okay? 3, 2, 1… So let's take the first one, tolerances. What does good sound like?
Terry Buechner: It may sound silly, but it's really thinking about it in a plain English, or whichever language, sentence that a business owner can understand, and a business owner can sign off on. For a particular workflow, we need to be back within a particular window of time with this level of data completeness, and with these controls still operating. It needs to be, if you think about it, board-readable, it needs to be testable, it needs to be objective, not subjective.
Neil Covington: And then the second one was test evidence. So you're saying not just a generic failover test.
Terry Buechner: And there is a place for those, too. I think, you know, regular generic failover tests are important, and technology and stress testing around that, too, is very important. But we need clear evidence around scenario-based testing that resembles real disruptions. And the cool thing is that with new tools and technologies, and we're not getting into a technology talk.
But there are a lot of ways for us to really model things, digital twins and other capabilities, to really model real disruptions. Late data, partial outages, access constraints, model run restarts, key person unavailable. all these things that… some of the things we talked about before, right, that are going to end up in those subtle changes that were… that we don't realize that have happened. So really… and trying big picture items, small ones, too. The point isn't that we're not getting to, like, we've recovered, we've turned things back on.
But, getting back to things we said before, that we've stayed within the tolerance guidelines that we have defined, and kept our controls intact. And if we find that that's impossible, or too stringent, or not stringent enough during these testing, then we can move different things in different directions around that, too. By doing these on a regular or scenario-based testing basis, then we can really have good, good model… good models, good guardrails, good controls in place. Tabletop exercise, very important, but only really if you're capturing decisions, making them tough, making them real, capturing decisions, capturing outcomes, capturing the remediation around that, too.
Neil Covington: And the third point was monitoring and attestation. Things change, things drift.
Terry Buechner: Exactly. Pipelines change, permissions shift, the models they update, teams reorganize, we lose people, bring people on, of course, all these things change over time. Big ones, acquiring a new business, divesting a business. So, the industry is moving from kind of this approach of annual DR testing, and it has been for a while, to be fair to the industry, we have been moving away from that for a while, both internally for our own purposes, but also because regulators have been pushing that, too.
But getting to this point of continuous posture monitoring, really be able to look at things in real time, test things in real time, have automated alerts as well. We don't necessarily need to throw tons of people at it on a regular basis, but to be able to surface this information on a regular basis, super important.
And we have the tools now to do it as well, which is a cool thing. Modernizing the data state, modernizing the core systems that are allowing that, very important to that, of course. We don't have the data available, we're not going to be able to have that continuous view. So we've evolved, the industry's evolved, and continues to evolve away from that periodic validation to that continuous approach around it. the… I think the platform's ongoing quality and availability metrics, they become really another piece of that resilience attestation. You can show that to the market, you can show that to your board, you can show that to regulators on a regular basis.
Neil Covington: And I'll just bring back one more element that you called out earlier, so the dated lineage as evidence as well.
Terry Buechner: Yeah, it's huge, right? I think that a governed data platform with that lineage tracking in it. Quality monitoring, versioning, that in and of itself, having that governed data map platform, modern data state, really creates a lot of your evidence. You can dive in and find stuff on the fly, you can have regular reports being produced, etc. You can show what's changed, what was impacted, when it was impacted.
And, if you remember Sherman and Peabody from, and their Way Wayback Machine, or Wayback Machine, the time travel ability around that, we've always been able to do that to a limited extent, but now querying data where it existed at a point in time. It really lets you reconstruct the decision path, what happened before the incident, where were we before the incident, what happened during the incident, what was used. was affected, what the outputs were downstream, who signed off, when they signed off, all those things, then we can really have that… the evidence-based approach to that, both looking back at what happened and also preparing for the future.
Neil Covington: For sure. So that moves resilience from, we think we're okay, to we can prove what happened, and what controls were operating. So let's go practical. Step by step, so without turning it into a multi-year program. We promised a practical path, in our chat. So, we've got five steps. So, Terry, do you want to start us off with step one?
Terry Buechner: Sure. Step one. the key… they're all very important, of course, but I think mapping those workflows, those actuarial workflows, and the dependencies around them. It's, again, not about the servers, not about the IT, those things are support, but it's about the data, it's about the business goals behind that. Where the data's ingested, how it's transformed, the model runs themselves, the approvals and sign-offs, all the downstream reporting. identifying single-source feeds, any silent enrichment steps, there's… many of those happen in the complexities of an insurance enterprise. there are a lot of enrichment that happen along the way, happen in an automated way in many cases, that people forgot about, that were never documented, or that teams left, or the systems have changed, right? So there's a lot of steps within that that need to be documented, and going back and looking at that.
there are many… this may seem like a big effort, I think, but it is hugely important, both for the specific areas we're talking about here, the actual workflows, but thinking about that from an overall workflow and efficiency perspective. So there's a lot of benefits that come out of going through this process. Mapping it out, you discover, oftentimes, I've seen where what seems like a single system risk, there's actually, and this not… should not be really surprising to anybody, but it's actually portfolio lever exposure, or at least much bigger than we thought it was.
So, one delayed feat, where we thought originally that this is only going to affect one single system or workflow, really it's affecting numbers across multiple lines of business, and really provide… can create a portfolio-level risk. That's why we need to look at a workflow level, and really then take out… move out to that system level view. But just having the system level view is not enough.
Neil Covington: And importantly, pick workflows that create real exposure. So if they miss decision windows, so for example, capital, reserving sign-off, stress scenario cycles, and so on, that's where you start.
Terry Buechner: Absolutely. And setting the impact tolerances by workflow, again, not by system. the… think about… we can think about data products, right? Like a curated, governed data set that's published for a specific business use, for consumption by business users. as our… think about those as our units of resiliency around that. I think when we get down to that level, then we really start to understand the business workflows, too. And there are knock-on effects around that beyond risk, right? Just efficiency or understanding how data is passed through and decisions are made throughout the enterprise. But you can think of maybe a… the statement could be an active claims data product must be available within 3 days after disruption, or 3 hours, whatever happens to be, whatever that SLA is appropriate for that. But it's business meaningful, which is super important, and testable, which is also very important. Again, it's very… it's an objective, not subjective approach.
Neil Covington: So workflow, not systems. And Step 3 is the one people tend to postpone into a crisis, so let's bring it forward.
Terry Buechner: Yeah, so what does it look like to be degraded and controlled in those operating modes in a crisis? It's something that… Also, perhaps, people are uncomfortable with this idea of running in a degraded mode. No, we have to run at 100% all the time, and that creates a tendency toward looking at a system view of uptime versus not time. Getting back up means that everything is working 100%. But as we've talked about.
There are situations where that is not acceptable. You need to have something signed off on. And so we need to define what degraded but controlled looks like. And perhaps that requires a conversation with regulators in some cases, too, or the auditors and the board, certainly, perhaps. But defining that, what does degraded but controlled look like? What can we proceed with documented caveats, documented sign-offs? What has to stop until we're fully back up? Who are the decision makers in that? encoding this, really, into your decision quality rules is super important. We can't be making anarchist judgments. That's gonna happen from time to time, of course, but we want to avoid that as much as possible. Avoid that idea as freelancing, coming back to our fire department strategy, right? Our analogy, excuse me. Doing that often creates… it seems like a good idea at the time, but often creates more risk than the actual disruption itself.
Neil Covington: And that's also where the governance conversation really starts, and really starts to get real. The workaround you choose must preserve defensibility, your approvals, your audit trail, and your data checks.
Terry Buechner: Absolutely, and I think that brings us to step four, which is test, test, test, test, having evidence, test against those tolerances. Tabletop exercises, around governance and decision making, but tie that back to technical exercises, too. You know, we've avoided talking about because we're talking about the business side of things, but clearly the technical pieces are very important.
But they can't be done in complete silos, and this is a tendency that we have across a variety of areas within insurance. to operate in silos, to think in silos, to invest in silos, to test in silos, really needs to be across the board at an enterprise level. We can, of course, have specific areas, and that makes sense, too. Like, focus on a specific workflow, or line of business. a specific set of technical aspects of it, too. That's important as well, but we have to think about it from a holistic, enterprise-wide, system-wide approach. data lineage, quality evidence, using those as test artifacts, what happened in the tabletop exercise, or what happened in the real disaster as well. Track that, understand it, be able to audit it, track that remediation, and again, plan for the future.
Neil Covington: And step five, make sure it sticks, doesn't it?
Terry Buechner: Yeah, that's often a very hard part, too, is not to… I don't want to criticize our industry too much. We have moved away from the once annual TR testing. Putting it and making it more continuous, approaching a continuous approach around that as well, but put it into business as usual. ownership, cadence, KPIs, board reporting, it really needs to be a living discipline. It's not a one-and-done project. We need to bake this into our business as usual.
Neil Covington: Exactly. And thought for the audience, so if you did just one thing next week, pick one workflow, reserving, capital modeling, etc. Write down, number one, the decision deadline. Number two, the impact tolerance. 3. You degraded but controlled ruled. And four, what evidence you show afterwards. Let's land this with a few questions that the audience can take away.
So, these are the ones to take back to your actuarial committee, or your risk governance forum. Which single actuaro workflow would create the most board exposure if it missed its decision window? What evidence could you show tomorrow that you can recover, and keep controls intact? And if a data feed into, for example, your reserving process was silently delivering stale data for 48 hours, how would you know?
And what would be the blast radius? I think that's a great place to stop. So, thank you, Terry, for your time and your contribution today. Ultimately, resilience isn't about keeping systems running, it's about maintaining confidence in the decisions those systems produce. Organisations that will lead are the ones that can demonstrate, with clarity and evidence, that their critical risk outputs remain trustworthy, even under disruption. Thank you, everybody, for joining us today. Discover more about the FIS Insurance Risk Management Suite - https://www.fisglobal.com/products/fis-insurance-risk-suite.Hope you found it useful, and feel free to reach out to myself or Terry if anything sparked interest.
Thank you very much.
Why is that the right framing right now?
Terry Buechner: Thanks, Neil, and first of all, thank you very much for having me. Very excited for the talk that we're going to have around this. It's very cool the way that it's been framed, I think. I think maybe interesting or surprising to have somebody from AWS and not talking specifically about IT or AI or other things.
And that's very important, because we're talking about it from all this resiliency risk from a business perspective, which is, as you said, how our teams actually experience it, right? And I think the cool thing is that, or interesting, is that resilience for those teams, for actuarial teams, risk teams, it's really a… you can think about it as a credibility issue. It's not an IT uptime metric. And when I say credibility, I'm talking about the numbers at the end, these key, important numbers, capital, reserves, stress scenario outputs, things like that.
And if you think about it, why it's important not to think about it purely in the sense of IT, although that is an important piece of it, of course, it's because how we came to these numbers at the end, it's from a long data supply chain, and it is truly supply chain. It's not a physical product in this case, but it is a supply chain, and this supply chain has been built by our carrier customers over, in some cases, decades, right? If you think about all the policy administration engines, administration systems, excuse me, that are in there, which are… can be manifold, at least more than a few for more, most carriers.
Claim systems, external data, and it's super important to be ingesting the vast amounts of available and more of data becoming available every day. If we're not taking advantage of that data and not ingesting it in a correct way, then we're not getting a full picture of risk overall. But when you get that data, whether it's internal or external, you need to enrich it, think about ETL pipelines, all of that.
And as, as I said, we built… we had carriers have built these chains, these, supply chains, incrementally over decades for them, right? So that… and it's through acquisition, through buying, through building. So when disruptions happen, and they do happen, whether it's a major one or a minor one. Often, although the disruption itself is certainly a big issue, and there's problems around that, but it's often the subtly wrong results, right, things that you don't know are wrong, or slightly off, and they're not obvious, that can be caused by partial or stale inputs that cannot be seen until almost it's too late, in many cases.
Neil Covington: Yeah, I think that's the key point, Terry. Wrong and unnoticed is worse than down and obvious. And we're seeing this becoming more than best practice, more like expectation.
Terry Buechner: Yeah, I think from the expectation there, and in insurance, it's interesting, I think that because we're… although there's a lot of interesting technology, and we've been using technology for a long time, almost since the beginning of the insurance industry. it is… we're a bit slow moving in these things, and a lot of the pressure comes from outside of the insurance industry. So, when we think about how regulators, auditors, shareholders, boards, policyholders are expecting us to not just say that Something is true, but rather demonstrate these things, often in real time, or more in real time as well.
So demonstrating operational resilience is very important, and you think about what regulators are expecting us to be able to provide to that, the bar is getting set higher and higher all the time. The other side of that is that the tools are available for us to do so. And part of that is modernizing the data state, right? So we see many insurers modernizing their core systems, modernizing their data states. Moving from a fragmented data state to a more unified platform, and that's for a variety of reasons. It can be business-driven, meaning growth, or profitable growth in terms of modernization, business agility. So, all those things tied together, and tied together very nicely if we think of them outside of individual silos. So, it's a modernization move, yes.
It's an IT resilience move, yes, but it is a overall risk resilience move. If you think about having fewer point-to-point integrations. a clearer data lineage. So when something breaks, well, first of all, you have fewer vulnerable points throughout, if you have a unified approach and uniform workflows. But you also have a clear view of what broke, when it broke, and be able to go back into a point in time and try to fix that.
Neil Covington: It sounds like the transition is, if resilience is a risk control, not a checkbox, what does that actually mean in actuarial terms? So if we break that down, using the triad we've been using, continuity, recovery, and control, perhaps, first continuity. What should actuaro leaders be doing differently, do we think?
Terry Buechner: I like this triad approach, having sort of a very clear, mechanisms around that. We're very big at mechanisms at Amazon, so I think having mechanisms around that and definitions is very important. So if you think about continuity. I like to think about it, and I've seen carriers successfully think about it in terms of decision windows. So, if you think about… and not the systems, although, again, important to remember that we're not saying the systems aren't important, they are very important.
But thinking about this from the business perspective. So, from continuity perspective, we think, what workflows, which ones of my workflows must still run, no matter what? Maybe it's reserving sign-offs, or calculations. you know, there are deadlines, hard deadlines in many cases associated with these ones. So, identifying those. And thinking about what those windows are is key to that. Some other analytics could be business intelligence or other things, maybe they can pause. So defining the criticality of them and the defining around that is super important around this, and not which server is the most important one to bring up, because it has our core policy system on it. Rather, what is the… what is the decision deadline that we have to sign off on something and think about it in those terms?
Neil Covington: So it's more workflow-based, not infrastructure-based, would you say?
Terry Buechner: Yeah, absolutely. So, it might be funny to hear from a cloud infrastructure provider to say that. The infrastructure is very important, of course, because it's the systems that support all these things. But we need to work backwards from what we're trying to achieve from a business perspective. The workflows are the key here.
And I think the data integrity is that, as we talked about a little bit, the hidden resiliency risk, right? So the data integrity, if a feed is delayed, or maybe it's partial, maybe the data is stale. So, perhaps the platform can still run, because there wasn't an obvious error in it. So the platform… so the workflow is still running. But the output maybe is misleading at that point. So the risk isn't the downtime itself, although that is a risk, but it's also that, what we would call, and we talked about this as the idea of these silent error propagation, something you don't know that's happening and can continue and have knock-on effects downstream.
Neil Covington: Okay, so if we think about a concrete example, perhaps, so say it's the morning of reserving sign-off, and a key claims feed is only partially in. What does resilient look like?
Terry Buechner: I like to think about it in terms of having a predefined, degraded, but controlled mode, and you can call that different things, but the data is degraded, the workflow is degraded, but we're still in control of it, we understand what's happening. So, do we decide, again, based on those, on those decision windows, do we decide that we run on the last good snapshot? Do we apply a conservative adjustment to it, or do we have to pause and wait until it's fully back up? Very important to that is who has the authority to decide, that. We need to identify that, have it very clear, have backups to that, too. Perhaps somebody's on vacation. We need to be able to capture that, capture those signs off… sign-offs, make sure it's very transparent, so we need to be able to defend to auditors, to the board, to regulators, why a decision was made at a certain point of time.
Maybe it's, do you know which reserve segments are affected by the missing data? Can you produce a result for the segments that you do trust? Flag the affected segments, etc. There's a lot of pieces around that. Get conditional sign-off, maybe. Degraded but controlled. So we understand what's going on, have very… we're clear about the decisions we made around that, we have control over it, not just run it and hope for the best, and hope for that things will work out at the end.
Neil Covington: And people will typically hear the word recovery, and what they actually hear is RTO and RPO, and immediately think IT.
Terry Buechner: And that's understandable, and I think the interesting thing is that in the… if you go back in time, sure, it began with a business process, thinking about RTOPO from a business perspective, but the underlying IT supports that, and over time, in my opinion anyway, is that it's become more of an IT issue rather than… we've kind of lost the original business decisions that led to that.
So we need to think about it in business terms. We're going to resurface that, really think about that in business terms, and understand the IT that supports that as well. Again, those decision windows, the reporting cutoffs, any tolerance for delay that there might be, whether that's internal, from a risk perspective, external to our customers, or the regulators around that, too. So it's not… restore the database in X hours, but rather, what's the latest point that we can resume, but still make that critical decision on time? Or how much can we delay that and tolerate before the output becomes misleading? Or can we pause this because it's not a critical workflow that needs to be brought forward until we're fully back up and running?
Neil Covington: And then the third leg, the control, that's one that's easy to overlook, I see.
Terry Buechner: easy to overlook in a critical situation. It happens to everybody in all sorts of situations, but when you're in that disruptive period, which of course always happens at, like, 2, 3 in the morning, you're under pressure. So there… there could be… and that pressure could be internal, external, a variety of things, but get that run out, we need to get sign-off.
But the control around it, access, change control, the data checks, approvals, audit trail, etc, all that is essential. And almost more importantly during these times of crisis, they can't be optional, right? Again, you can't just get the run out, you can't hope for the best. you know, you need them the most at these points. So that resilience needs that you have, the degraded but controlled situation going on.
And be able to control that and understand what's going on. And Neil, I think you know this by now, but I'm a volunteer firefighter, been doing that for about 12 years now in my hometown, and there's this concept that is applicable here, it's called… we call it freelancing, which means that you go off and do your own thing on the fire ground, in an emergency situation. Because either it seems like a good idea. Or, more often, and more dangerously, because you're in that crisis mode, something's happened. Something's happened, and you make a snap decision, and react on instinct. and hope for the best, rather than having that control of training and chain of command. And on the fire ground, that can, you know, in a worst-case scenario, it can get people hurt, right? But in best case, it's going to be self-optimal as it comes out.
In an actuarial workflow, those ad hoc workarounds, which are used very, very often during disruptions, they bypass the controls that give you, and give our… give the auditors, good regulators, etc, confidence in the output. And you may not discover that damage until it's too late, and you may not see that damage until way down the road, like during an audit.
So it might not be physically dangerous, like, on the fire ground, but it still can be dangerous to the enterprise, for sure.
Neil Covington: And kudos for being a volunteer firefighter, Terry, for sure. I mean, I wonder how many of the audience have an explicitly documented degraded mode for reserving or capital flow, capital workflows and so on, including who signs off the exceptions.
Terry Buechner: Very good question, and I think that needs to be looked at on a regular basis, you know, having the resiliency more than just once a year, thinking about DR, all those kinds of things, and I do think that That modern data platform approach can really help Making it more explicit, whether you're dealing with that missing data means that we stop, we pause, versus stale data means we can proceed, but we have some reservations or caveats around that, too. It's a control. It's a control that we have, and we have documented it's not an opinion formed at 2 in the morning.
Neil Covington: Absolutely. And a perfect transition, so… Having controls is necessary, but boards, auditors, regulators will ask the same follow-up, prove it. And from my experience, the… I've always been told, in terms of audit, if you can't prove it, it hasn't happened. Learning point two, it's perhaps the difference between believing your resilience and proving it, then.
So, we've been using a blunt line. If your only evidence is a DR slide deck, you have an assumption, not a control. Why does that resonate, do you think, Terry?
Terry Buechner: I think that we need to have that in place because different audiences are going to expect different, proofs, right? They're going to… the resilience expectations are shifting toward, excuse me, toward evidence, and the different audience want different types of evidence behind that. Kind of different motivations, if you will. And there are many different stakeholders, but some of the key ones I think about are, of course, the board. They want impact tolerances. What outcomes are at risk? What are the tolerance thresholds for the enterprise as a whole? Auditors, they want to see control design, they want to see operating effici- effectiveness.
Our regulators, of course, they want tested ability to stay within whatever the tolerances they've defined, and there are many different definitions of that, many different regulatory bodies and different approaches to that, too. Very important to remember, we have shareholders in many cases, not all cases, of course, with mutual, etc, but shareholders, they want that steady ship. They don't want to see chaos, they don't want to see that ship tossed into chaos, into the stormcast during a disruption. Or, perhaps even worse, being pinged by regulators because of that chaos.
And then policyholders, which can be in a mutual, the shareholders, too, but the policyholders, really, they just, for the most part, they want assurance that their claims are going to be covered, that no matter what, their claims are going to be covered in a timely fashion. So they… and they want evidence of this, too. That evidence for policyholders might be lack of being pinged by a regulator, but still, that is an important thing to that. So having that evidence is important and needs to be defined by the audience as well.
Neil Covington: So how do we convert resilience into things you can actually show?
Terry Buechner: Think that, you can think about it in terms of, Proof artifacts, if you will. Sorry, there was a… I thought I had turned off my announcements, but that came up, so let me just start back from, neil, can you ask that question again?
Neil Covington: Okay. I'll just pause for a moment. So, 3, 2, 1… So how do we convert resilience into things you can actually show?
Terry Buechner: Approach can be to have three proof artifacts, if you will. One is having documented tolerances, documentation is very important, and those key recovery objectives, but again, not from an IT perspective, but from a key workflow perspective. One is having the test evidence behind it, as we talked about a bit, scenario-based And one is having the ongoing monitoring and attestation. Make sure that the posture hasn't drifted, but be able to test and visualize these things on a regular basis, too.
Neil Covington: So let's take the first one, so tolerances. What does good sound like?
Terry Buechner: Great, let's try that again. Another announcement came up that I tried to turn off, I apologize, please. But I've turned them all off now, so please ask that question. Let me just do one more thing, sorry. It's off. Should be good. I apologize, please continue.
Neil Covington: That's alright. You could hurry up. Okay? 3, 2, 1… So let's take the first one, tolerances. What does good sound like?
Terry Buechner: It may sound silly, but it's really thinking about it in a plain English, or whichever language, sentence that a business owner can understand, and a business owner can sign off on. For a particular workflow, we need to be back within a particular window of time with this level of data completeness, and with these controls still operating. It needs to be, if you think about it, board-readable, it needs to be testable, it needs to be objective, not subjective.
Neil Covington: And then the second one was test evidence. So you're saying not just a generic failover test.
Terry Buechner: And there is a place for those, too. I think, you know, regular generic failover tests are important, and technology and stress testing around that, too, is very important. But we need clear evidence around scenario-based testing that resembles real disruptions. And the cool thing is that with new tools and technologies, and we're not getting into a technology talk.
But there are a lot of ways for us to really model things, digital twins and other capabilities, to really model real disruptions. Late data, partial outages, access constraints, model run restarts, key person unavailable. all these things that… some of the things we talked about before, right, that are going to end up in those subtle changes that were… that we don't realize that have happened. So really… and trying big picture items, small ones, too. The point isn't that we're not getting to, like, we've recovered, we've turned things back on.
But, getting back to things we said before, that we've stayed within the tolerance guidelines that we have defined, and kept our controls intact. And if we find that that's impossible, or too stringent, or not stringent enough during these testing, then we can move different things in different directions around that, too. By doing these on a regular or scenario-based testing basis, then we can really have good, good model… good models, good guardrails, good controls in place. Tabletop exercise, very important, but only really if you're capturing decisions, making them tough, making them real, capturing decisions, capturing outcomes, capturing the remediation around that, too.
Neil Covington: And the third point was monitoring and attestation. Things change, things drift.
Terry Buechner: Exactly. Pipelines change, permissions shift, the models they update, teams reorganize, we lose people, bring people on, of course, all these things change over time. Big ones, acquiring a new business, divesting a business. So, the industry is moving from kind of this approach of annual DR testing, and it has been for a while, to be fair to the industry, we have been moving away from that for a while, both internally for our own purposes, but also because regulators have been pushing that, too.
But getting to this point of continuous posture monitoring, really be able to look at things in real time, test things in real time, have automated alerts as well. We don't necessarily need to throw tons of people at it on a regular basis, but to be able to surface this information on a regular basis, super important.
And we have the tools now to do it as well, which is a cool thing. Modernizing the data state, modernizing the core systems that are allowing that, very important to that, of course. We don't have the data available, we're not going to be able to have that continuous view. So we've evolved, the industry's evolved, and continues to evolve away from that periodic validation to that continuous approach around it. the… I think the platform's ongoing quality and availability metrics, they become really another piece of that resilience attestation. You can show that to the market, you can show that to your board, you can show that to regulators on a regular basis.
Neil Covington: And I'll just bring back one more element that you called out earlier, so the dated lineage as evidence as well.
Terry Buechner: Yeah, it's huge, right? I think that a governed data platform with that lineage tracking in it. Quality monitoring, versioning, that in and of itself, having that governed data map platform, modern data state, really creates a lot of your evidence. You can dive in and find stuff on the fly, you can have regular reports being produced, etc. You can show what's changed, what was impacted, when it was impacted.
And, if you remember Sherman and Peabody from, and their Way Wayback Machine, or Wayback Machine, the time travel ability around that, we've always been able to do that to a limited extent, but now querying data where it existed at a point in time. It really lets you reconstruct the decision path, what happened before the incident, where were we before the incident, what happened during the incident, what was used. was affected, what the outputs were downstream, who signed off, when they signed off, all those things, then we can really have that… the evidence-based approach to that, both looking back at what happened and also preparing for the future.
Neil Covington: For sure. So that moves resilience from, we think we're okay, to we can prove what happened, and what controls were operating. So let's go practical. Step by step, so without turning it into a multi-year program. We promised a practical path, in our chat. So, we've got five steps. So, Terry, do you want to start us off with step one?
Terry Buechner: Sure. Step one. the key… they're all very important, of course, but I think mapping those workflows, those actuarial workflows, and the dependencies around them. It's, again, not about the servers, not about the IT, those things are support, but it's about the data, it's about the business goals behind that. Where the data's ingested, how it's transformed, the model runs themselves, the approvals and sign-offs, all the downstream reporting. identifying single-source feeds, any silent enrichment steps, there's… many of those happen in the complexities of an insurance enterprise. there are a lot of enrichment that happen along the way, happen in an automated way in many cases, that people forgot about, that were never documented, or that teams left, or the systems have changed, right? So there's a lot of steps within that that need to be documented, and going back and looking at that.
there are many… this may seem like a big effort, I think, but it is hugely important, both for the specific areas we're talking about here, the actual workflows, but thinking about that from an overall workflow and efficiency perspective. So there's a lot of benefits that come out of going through this process. Mapping it out, you discover, oftentimes, I've seen where what seems like a single system risk, there's actually, and this not… should not be really surprising to anybody, but it's actually portfolio lever exposure, or at least much bigger than we thought it was.
So, one delayed feat, where we thought originally that this is only going to affect one single system or workflow, really it's affecting numbers across multiple lines of business, and really provide… can create a portfolio-level risk. That's why we need to look at a workflow level, and really then take out… move out to that system level view. But just having the system level view is not enough.
Neil Covington: And importantly, pick workflows that create real exposure. So if they miss decision windows, so for example, capital, reserving sign-off, stress scenario cycles, and so on, that's where you start.
Terry Buechner: Absolutely. And setting the impact tolerances by workflow, again, not by system. the… think about… we can think about data products, right? Like a curated, governed data set that's published for a specific business use, for consumption by business users. as our… think about those as our units of resiliency around that. I think when we get down to that level, then we really start to understand the business workflows, too. And there are knock-on effects around that beyond risk, right? Just efficiency or understanding how data is passed through and decisions are made throughout the enterprise. But you can think of maybe a… the statement could be an active claims data product must be available within 3 days after disruption, or 3 hours, whatever happens to be, whatever that SLA is appropriate for that. But it's business meaningful, which is super important, and testable, which is also very important. Again, it's very… it's an objective, not subjective approach.
Neil Covington: So workflow, not systems. And Step 3 is the one people tend to postpone into a crisis, so let's bring it forward.
Terry Buechner: Yeah, so what does it look like to be degraded and controlled in those operating modes in a crisis? It's something that… Also, perhaps, people are uncomfortable with this idea of running in a degraded mode. No, we have to run at 100% all the time, and that creates a tendency toward looking at a system view of uptime versus not time. Getting back up means that everything is working 100%. But as we've talked about.
There are situations where that is not acceptable. You need to have something signed off on. And so we need to define what degraded but controlled looks like. And perhaps that requires a conversation with regulators in some cases, too, or the auditors and the board, certainly, perhaps. But defining that, what does degraded but controlled look like? What can we proceed with documented caveats, documented sign-offs? What has to stop until we're fully back up? Who are the decision makers in that? encoding this, really, into your decision quality rules is super important. We can't be making anarchist judgments. That's gonna happen from time to time, of course, but we want to avoid that as much as possible. Avoid that idea as freelancing, coming back to our fire department strategy, right? Our analogy, excuse me. Doing that often creates… it seems like a good idea at the time, but often creates more risk than the actual disruption itself.
Neil Covington: And that's also where the governance conversation really starts, and really starts to get real. The workaround you choose must preserve defensibility, your approvals, your audit trail, and your data checks.
Terry Buechner: Absolutely, and I think that brings us to step four, which is test, test, test, test, having evidence, test against those tolerances. Tabletop exercises, around governance and decision making, but tie that back to technical exercises, too. You know, we've avoided talking about because we're talking about the business side of things, but clearly the technical pieces are very important.
But they can't be done in complete silos, and this is a tendency that we have across a variety of areas within insurance. to operate in silos, to think in silos, to invest in silos, to test in silos, really needs to be across the board at an enterprise level. We can, of course, have specific areas, and that makes sense, too. Like, focus on a specific workflow, or line of business. a specific set of technical aspects of it, too. That's important as well, but we have to think about it from a holistic, enterprise-wide, system-wide approach. data lineage, quality evidence, using those as test artifacts, what happened in the tabletop exercise, or what happened in the real disaster as well. Track that, understand it, be able to audit it, track that remediation, and again, plan for the future.
Neil Covington: And step five, make sure it sticks, doesn't it?
Terry Buechner: Yeah, that's often a very hard part, too, is not to… I don't want to criticize our industry too much. We have moved away from the once annual TR testing. Putting it and making it more continuous, approaching a continuous approach around that as well, but put it into business as usual. ownership, cadence, KPIs, board reporting, it really needs to be a living discipline. It's not a one-and-done project. We need to bake this into our business as usual.
Neil Covington: Exactly. And thought for the audience, so if you did just one thing next week, pick one workflow, reserving, capital modeling, etc. Write down, number one, the decision deadline. Number two, the impact tolerance. 3. You degraded but controlled ruled. And four, what evidence you show afterwards. Let's land this with a few questions that the audience can take away.
So, these are the ones to take back to your actuarial committee, or your risk governance forum. Which single actuaro workflow would create the most board exposure if it missed its decision window? What evidence could you show tomorrow that you can recover, and keep controls intact? And if a data feed into, for example, your reserving process was silently delivering stale data for 48 hours, how would you know?
And what would be the blast radius? I think that's a great place to stop. So, thank you, Terry, for your time and your contribution today. Ultimately, resilience isn't about keeping systems running, it's about maintaining confidence in the decisions those systems produce. Organisations that will lead are the ones that can demonstrate, with clarity and evidence, that their critical risk outputs remain trustworthy, even under disruption. Thank you, everybody, for joining us today. Discover more about the FIS Insurance Risk Management Suite - https://www.fisglobal.com/products/fis-insurance-risk-suite.Hope you found it useful, and feel free to reach out to myself or Terry if anything sparked interest.
Thank you very much.