Summary
- Introduces DataConnect v14 with integrated data quality capabilities.
- Automates quality assessment and issue detection across pipelines.
- Improves visibility into data lineage and pipeline health.
- Previews AI-assisted features for quality management and integration.
For those joining, we're going to give it a minute or two, and then we'll get going on the Data Connect webinar. All right. Looks like we've got a couple more people coming in, so we'll just give it maybe another half minute, and then we'll kick this off. Okay, why don't we go ahead and get going? We're two minutes after. Welcome, welcome, guys. Thank you so much for making time to join us for our Data Connect webinar.
We're really excited to go over some of the new features and functionality within Data Connect. Before we do, there's a few housekeeping items that we're going to touch on really quick. So just as you probably already see, you're all muted. Feel free, though, to submit questions in the Q&A. Just to put an emphasis, the Q&A panel is on the More button, so not in the chat. We are monitoring the Q&A. This will be recorded, so we'll share the recording out along with any additional resources for you in follow-up to this.
And then we scheduled 45 minutes, so hoping to do questions there at the end as well, live if we can. But again, thanks for joining, and with that, we'll go ahead and jump right in. We'll first start with some quick introductions. So Chris, I'll pass it to you. Thanks, Lauren. Hi, guys. My name's Chris.
I'm the VP of Data Connect Engineering here at Actian. I've been with the product for just over 25 years. We've been working awfully hard in the engineering department, and I'm super excited to show you all the good things that we've brought to the table for Data Connect. We're going to get into our data quality release, which is V14, and the following release, which is Integration Manager 14.1. So back over to you, Lauren. Awesome. Thanks, Chris.
And I'm Lauren Cartwright, head of sales here for Americas at Actian. Been working really closely with Chris and team, as well as our customers, and just very, very excited to show you guys what we're doing within Data Connect, where we're going, and then just some of the internal changes we've made as an organization too that are really intentional around coming around our customers and driving better value-based outcomes. So with that, we'll go ahead and go to the next slide. And a little bit about what we're going to touch on today, right? We're hearing a lot around data hype versus data lifecycle. AI, we're in this age of AI, no different than big data years ago. As we continue to innovate, everything always comes back to how trusted is your data, and Data Connect has always been an integral part of that.
But we have gone ahead and made a number of enhancements to it, especially around quality, that is really enabling us to offer our customers better visibility, observability, quality, and trust across the entire lifecycle of your data. We're really excited about this because not only have we made these enhancements to provide that better visibility, but we've done it so that every user can use those capabilities, regardless what your expertise is in the product or in general around data management tools. In addition to that, as an organization, we've really put our focus around this. So starting in April of this year, I stepped in and have an entirely new team that is dedicated to supporting the data management function within Actian, and Data Connect is the connective tissue within that function. Everybody on our team comes from a data management background. I myself have deployed over a dozen data governance programs across multiple industries and multiple tools. And so we're really leaning into making sure that we are centered around those value-based outcomes for our customers.
And that includes making the enhancements, the very intentional enhancements, around Data Connect that also enable us to extend entitlement to you guys to make the upgrade path much easier. So we'll get into that a little bit, but all of that is, again, with that intention to really bring a better experience to our customers and those value-driven outcomes. By enabling that, we are making sure we lean into the stronger platform. So you're going to see the refreshed UI that is way more than just a refreshed UI. It is really about bringing it to everybody. You're going to see the AI-powered data quality capabilities that are going to take tasks that took months to now minutes, enhanced monitoring, and the easier upgrading for you to V14. So we're really excited to show you guys.
With that being said, we can go ahead and go to the next slide. And this just gives you a little bit of color on our roadmap. We're of course, more than happy to schedule one-on-one calls following this webinar and dig into more detail. But today, as we go over V14, you're going to see what we're doing in 14.0 with regards to the data quality and enriched data prep. This is really about that automation, that bringing it to every user, the speed to get value out of data quality so that you can really drive trusted outcomes, whether you're driving towards AI initiatives or you're doing migration projects. Really any data management strategy is going to center around that trusted data, and we've got some really cool things to show you guys. And then in Q3, we're making enhancements to Integration Manager, further enhancing monitoring, measuring, and managing capabilities within it.
So you're going to see some really, really rich observability. It gets me super excited. You can see it across the entire lifecycle of your data embedded there in the pipeline. So now your meantime to resolution, much, much easier to reduce that and solve for. Enriched lineage, AI-assisted analysis. Again, a lot of exciting things there. Q4 is all about advanced data integration.
So this is where you're going to see an easy upgrade path, forward/backward compatibility, improved map designer, process designer, schema designer. So again, really exciting enhancements. And then we get into the early start of next year, and that is really all about bringing everything that we've put into Data Connect to really any tool, right? Becoming truly tool-agnostic, the orchestration and observability across your entire ecosystem. Again, these capabilities, I think you're going to find are completely new to Data Connect in many ways, but again, really hinge on the core values of Data Connect that have long been there, that really drive that trusted usability of your data. So with that, I will turn it over to Chris, and again, thank you for being here, and don't hesitate to put questions in any part of the queue. Thanks, guys.
Thanks, Lauren. One second here while I share. There we go. Okay. So today, as Lauren mentioned, we're going to talk through data quality. I'm going to walk you through all the design options that you have there, all the ways that you can gain insight from our UI. The predominant themes within data quality is we wanted to accommodate all of our users by making data quality as simple as we possibly could.
And that meant giving you maximum insight into your data, being able to quickly identify the problems, and automate the rules needed to measure and resolve those problems. So we'll walk through all of that. We'll show you how we're using AI in the product. Let's see if I can get rid of that. Maybe I can figure it out later. I'll also show you Integration Manager, the next release, 14.1. As Lauren mentioned, we've got a lot going on in Integration Manager around observability, around lineage, even some design services within Integration Manager to help you maintain your jobs, keep them healthy.
So we're going to start right here with the Data Quality Editor that you're looking at right now. You can see that I've connected to a data set. I've got a view of my schema, a sample view of my data set here. And I'm going to click Inspect Data. And so what's going to happen when I'm clicking Inspect Data, we're doing two things. A, we're running a semantic mapping, and B, we're running data discovery. And so the semantic mapping helps us identify, for example, what kind of data is within that field or within any given field.
So, for example, we understand that it's a date field, and that helps us understand the appropriate rules that might be configured against that particular field, and also helps us understand a little bit about what's expected inside that field. The data discovery piece, which I'll show you in a minute, is going to help you understand the frequency of values and patterns and statistics within any given field. But first, the design options that you have here, automate design, design assistance, and manual design. So manual design just takes you straight into the editor. Automate design is where we're going to actually build the solution for you, and all you have to do is review it, make any changes that you need, and deploy it. Design assistance, to us, that indicates that you need help or want help in designing your solution. Maybe you're not that familiar with your data set, maybe you're new to data quality, and in this case, we're going to suggest rules to you, give you a very simple way of reviewing and validating those rules.
So with both of these two options, we have a menu associated with it. They're very similar, but in essence, all you need to do is tell us which fields you want us to process and evaluate. You have some optional columns over here that at the beginning of the design process, even if you're not familiar with it, this is where data discovery comes into play. I'm able to quickly see how many different date patterns I have for this date field. And, for example, where values are concerned, understand all the different values that I have in there and the frequency of those values. So if I wanted Data Connect to generate rules based on my preferences, all I have to do is come in and select one or more of these patterns, for example, one or more of these values, for example. And what we will do, we will take that and turn those into data quality rules, i.e., profiling and remediation.
You also have an easy way. For example, if I wanted to go and mask a first name, I have several choices here of the characters I could mask with. Within the product itself, I won't go through it during the demo, but we do have a de-identify rule where you have all sorts of different options in terms of the different characters you can use, the number of characters you want to mask, and you can also even encrypt if you wish. But what happens is when you start the scan, our engine's going to run, we're going to evaluate all the data, we're going to generate all the rules, and that's either going to be depending on which one you select. We're either suggesting the rules for you, or we're automating the entire thing. So let's take a look at an automated design real quick. So here you can see I'm in the editor.
This is our data browser. And over here on the right-hand side, I've got a list of all the different rules that have been generated against this data set. I have a single score up here called the Data Quality Index, and that represents trustworthiness for this data set. I don't have any dimensions associated with my rules. That is optional, and you can do that. And if you do that, we'll be generating all the metrics for those different dimensions as well. But there's two different ways that you can quickly validate.
One, you can start looking at each rule. And so if you click on this link right here, what it's going to do is expand, and you'll see how the rule was configured. You'll see what the rule is doing. And you can also drill down by clicking on this bar chart right here. So that bar is going to show me all the valid data. And in this particular case, there is no invalid data. But if there were, I'll show you plenty of examples in a moment.
All you have to do is click on it, and it's going to render in the browser for you. So that's one way of doing it, going rule by rule. But if you click on the Data Quality Index, we're repainting this panel, and we're reorganizing. So before we were showing you all the rules in order of the schema. Now we're showing them organized by profile and remediation. In the same way as before, I could go and review the rule, look at the clean data based on the rule. In this case, it is not blank.
I can look at the dirty data and see that in the browser as well. This view right here, this chart, we can visualize the data as an aggregate. So not in the context of the individual rule, but actually what's being written to our targets. So within Data Connect, you have two targets coming out of our engine, a pass target and a fail target. And how this works is your profiling rule contains a condition. So for example, we could have a condition around this date field specific to the format or to the value, right? And so that's what's inside these profiling rules, and that helps us enforce your preferences, what you consider to be valid.
It also helps us measure it too. So the remediation rule, what it does is it can take an invalid format or an invalid value and change it, cleanse it to where it becomes valid. And so when you're looking at these views of your overall pass, fail, and remediated, it's very similar to what I just showed you. I can click on the green portion here. Notice that the header is changing color. This is showing me all the valid data that's in my data set. So this is valid data that we didn't touch, we didn't remediate.
It was valid in its native state. If I click on the purple portion, the header will change to purple, so you can quickly identify that you're looking at remediated data. So this would be all the data that we've cleansed. And you can finally click on the red portion to look at all the dirty data. And so one of the things that I think is really powerful and valuable about the way that we've designed things is you can simply go through your profile rules by clicking on the red portions here to understand why that data is invalid. So if I'm clicking on this rule, which is configured against the date field, I can easily see I've got some invalid date formats. There's a character, an exclamation point in there.
That's an invalid date because there's no day. I can go and look at my emails, why are those invalid? Same thing, right? So kind of quickly getting to the point and quickly understanding what your problems are. I think that call out too, just to double down. To me, this is one of the most exciting features because it's holistic view, right? You can very quickly, like you just went through.
The date, the email. We're not toggling from one screen to the other, and I can reference what has been cleansed, what has failed, what was already clean. Again, all staying on the same page and same experience. So just to double down on that. Yeah, 100% agree with you, Lauren. And the other thing is, true, you don't have to worry about managing the data. We're already isolating it for you, right?
Those two separate targets. And as you go, and I'm going to show you in just a moment, while you're designing, as you're designing, you're able to quickly delineate between what's valid, what's invalid, and what's been remediated. So next, let's go take a look at the design assistant path. So this would be that middle tile that you saw on the scan menu. The first thing that's going to happen, because you're asking us for help, we're going to give you an assessment of what we found. And so up at the top, it's just reiterating what you asked us to do. In the middle, we'll tell you the different kinds of data quality problems we found along with definitions.
And then down at the bottom, we're going to map all those issues to your schema so you can see for any given field what kind of problem we found and how many records were impacted, right? And you're doing all that before you ever hit the editor. So going into the editor, you have a better understanding of what you're up against and what you need to do in your solution. So once we come into the editor, notice that we don't have any rules, so we don't have an index yet. There's three interactions here that I really want to draw your attention to. First is the browser. So if you've been using Data Connect, you've been staring at our browser for years, and I'm really excited to tell you we've made some changes to where your browser is now an editor.
And it functions in a lot of different ways. But first, what I'll show you, let's go pick a relevant field. Let's say phone number. So let's say I'm looking in my browser, and I've done this myself a million times over the years using Data Connect. I find something in the browser that just doesn't look quite right to me. So what we've done is we've allowed you to click it, and when you click it, notice that this refreshes to show me the rules that have been recommended to me. Again, I'm going down the design assist path.
I need help. And we're also highlighting that pattern down in data discovery. So again, quickly bringing insight to the user. So I understand now using the display pattern, what it's associated with and how many instances I have of that in my data set. And I can also easily see the rules that are being recommended to me, right? And there's more to come on the editor I'm going to show you in just a minute, but again, I'm super excited to show you. But going down the path of design assistance, we've scanned your data, and we've got this curated set of rules for you to validate, to review.
And so we've made that as simple as possible because this right here is the name of the rule. So it's the field name underscore the rule name, and all you have to do is click it. When you click it, what's going to happen is the browser's going to refresh. It's going to show you the invalid data per this rule. I can see how this rule is configured, and I can change it if I wish, but I can also see what's going on in data discovery in terms of the frequency of the pattern, the frequency of the values, the statistic, excuse me, statistics. And I can also get examples of valid and invalid data per this rule. And so I can also drill down on the valid data as well, right?
So inactive, active. I can see all that there. And looking at the same, I can see what my problem is, right? So if I like the rule, all I have to do is apply it. And so when you apply the rule, automatically our engine is going to fire, right? So now I have a data quality index, and I can quickly validate what I just did by clicking on this bar. Again, I'm looking at the valid data, and I'm looking at the invalid data, right?
To make it really, really easy to see what's going on within my design. Do I like the way this rule is configured? Is it behaving in the way I expect it to? So coming back to what I just said a moment ago about our browser now being an editor This is one of the many ways, which I'll show you, in how you can automate a rule within Data Connect. We still have the manual path if I wanted to go add a rule manually, and as I mentioned before, we're generating all the regular expressions for you should you ever need one. But one of the things we really wanted to do with Data Connect 14 in terms of data quality is make it to where you're not fumbling around through some long list of rules trying to figure out which is the right rule I need to use, and then spending time trying to figure out how to write a regular expression for what you're looking for. So you can see here I've got my list of dirty data per this rule, and I'm going to leverage data discovery to help me understand what's wrong with each of them.
This first one's pretty obvious what's wrong with it. We know that active and inactive are our only allowable values. But down here, this looks like that should be valid, so I don't quite get it. So if I just click on it, looking at data discovery, looking at the display pattern, that S right there tells me there's a leading white space there. Right? So that's our problem. And I can see with active and inactive, I've got a casing problem.
So if I single click the cell, that's going to cause this to refresh and cause data discovery to refresh. But if I double click the cell, it's going to go into edit mode. And so I'm just going to go and change the casing to active, and I'll commit my change. And when I do, I'm prompted on how I want this change to be interpreted. Do I only want to apply this change to the cell? I want to apply it to all cells with identical values, or it's a format change, leave the value alone. So I'm going to select format change and apply it.
And again, anytime you do anything with a rule, our engine's going to fire to make sure the entire thing, all these different components are all up to date with what you're doing in your design. Right? So it added the rule for me. I can go drill down on it to go validate quickly once again what I did to make sure it's working the way I expect it to. And within our browser, you're seeing the old value and the new value, right? The old value is crossed out in red, the new value is in green. That only appears in our editor.
That doesn't get written to the target. But I can easily go back to my profiling rule and look at my failed records again. Again, we're giving you a distinct list of your failures to help you focus and help you see what's wrong using that with data discovery. And I can see that my data quality index has improved. I also can see within data discovery that the number of patterns I have has diminished. Right? It's current with what I've done.
While we're on the topic of data discovery, I want to show you a little bit of what we've done there. So in this example, we're going to take a look at this renewal date field. So a cursory glance at the editor will show that we've got a lot of different patterns going on here. So I could either click in the editor or the browser itself to cause this to come up, or I could just select it from the menu as I just did. So looking at this, I can immediately see I've got seven distinct patterns within my date field. What's really cool here is I can click on any one of the entries in this table and cause the browser to refresh so I can see them clearly. And when I've made a decision on which I'd like to standardize on, I can click this button down here.
And at this point, all I have to do is select one or more patterns that I want to standardize on. I save it. I can see how the rule is going to be configured. I can preview my rule, and because we're making a format change, we're going to automate the generation of a profiling and a remediation rule for you. You can dismiss either of them if you wish, but just by clicking on these tiles, I can see how the rule is going to be configured. In this case, the remediation rule is going to convert all these different date formats to this format. All I have to do is hit apply rule.
Again, our engine will fire. I can see that my rules have been added, and now I can drill down and validate how they behaved and how they performed. And also notice our rule was able to catch, while the format was valid here, the date itself, the value is invalid. And again, I can go explore my remediation and see what the previous examples and notice all the different kind of date formats that we were able to convert very easily, right? So part of all this is, again, we're trying to bring focus to the data quality issues that you have, make it really easy to see it, see the trends, and be able to quickly automate the generation of the rules. I think also something to highlight is it's showing you where your various formats are before you're ever even doing anything, right? So if you notice you have a format that is associated with 22 different assets as opposed to the other formats that may only have one or two association, it may help organizations more quickly and readily prioritize which one they want to default to.
Oftentimes, that's what we kind of struggle with is, well, what is our standard? Or maybe we have two of five standards we use more than the others. This is very quickly telling you, again in a single view, where you may want to prioritize or serve as a default. Agree. Agreed. We've got so much to cover, so I want to go back to the scan menu really quickly, and show you a couple of things. So within the design assistance menu, we give you the ability to go and search for fuzzy match duplicate records.
So I know duplicate records is the bane of everyone's existence. Everybody hates them. Everybody has them. We make it really easy for you to go find fuzzy match duplicate records and also exact match duplicate records. So if you want us to build the design for you, all you have to do is check that button, and we will go and look for exact match duplicate records. I've got an example for you right here. You'll see, let me generate the preview here.
You'll see what we're doing. We have the ability within the automated design path to go identify and find all those exact match duplicate records. Since they're all identical, truthfully, it doesn't matter which one we keep, but our default behavior is we're going to keep one of them and write the rest to our fail file as duplicates. But in reality, you could come in here and select as many of these as you want if you wanted to keep them. That's kind of how exact match duplicates work. With the fuzzy match duplicate behavior All you have to do, it looks like I messed this up, but all you have to do is select a key field and then select one or more matching fields, and I'll show you how that works really quickly. I think I need to fix that bug.
So I'm going to show you how this works, and I'll show you an example of it. Effectively, you pick a key field, and within that key field, we're going to look for exact matches, and then you can pick as many matching fields as you want and select from this list of algorithms that we provide. And so how it works, if we found 10 exact matches for account number, we will create a cluster, and I'll show you exactly what this looks like in a second. Within that cluster, we're going to go look at these two fields and go find fuzzy matches based on this algorithm. Okay? And so what that looks like is we're going to create a cluster file for you to review, and this is what it looks like. And so this is our key field, first name.
And so you can see exact matches, and you can see the unique cluster ID that we've created for you. So all you have to do is come in and look at each cluster. Those are the exact matches. You can see how we fuzzy matched here and how we've done that across all these different clusters. All you have to do is come in and select which records you consider to be unique out of each cluster. We will write those to our pass file, our pass target, and the rest will be written to the fail target as fuzzy match duplicates. You're more than welcome to pick all those records in each cluster if you want to.
Multiple or single, whatever you need, whatever you like. Okay. Next, let's get on to our AI prompt. So within AI, within our offering here within version 14, you have the ability to plug in any AI service that you want. You also have the ability to pick any model that you want to, and you can even do it within the same session. If you want to be cost-conscious, you can pick a particular model. If you're doing something rather complex, you can pick a different model.
But how it works is our AI prompt shows up down here, and you can type hello, and it's going to respond with a list of all the different things it can do. I personally have used it. I've used it to filter data in our browser. If I'm looking for something and I want to see a subset of records based on some condition, I use it for that. I use it to generate, and automate profiling and remediation rules for me. I use it to debug. I use it to assess my data quality design and my solution and find gaps for me.
So it's very powerful. And right down here, as I mentioned, you can choose from your provider and also choose from the model that you want to use. And so right here, you can also clear your context. And so one of the things that I think is great about this prompt, it's one thing to be able to look at this revenue field and look at this cell and go, okay, the format is valid, and maybe it's within the appropriate range, so the value seems to be valid. But when you look at the business logic and the relationships within the dataset, being able to validate at that deeper level, I think is just amazing. And this AI prompt makes it super easy to do that. So within this dataset, I've got a field called tax rate, and I've got another field that has the actual state in it.
So I'm going to use this AI prompt, and I'm saying, make sure the tax rate field is correct for the state field for US states. And that's all I'm telling it. And so just as the same experience, if you're using Claude, for example, in your IDE, we're waiting on the model. It's communicating with the model. But what it's going to do is it's going to come back with a rule that it's suggesting to me, and I have the ability to allow it or skip it or modify it or do anything that I want to with it. But you can see it's come back, and it's told me, "Okay, I've created this rule. Is this what you want?" And I say apply.
And just like anything else I've shown you today, our engine is going to execute. The rule is going to get added. And once it's added, I have the ability to go drill down on it and look at the pass data and the fail data. And by the way, we've also added this, and this is specific to the browser. This doesn't change how the data is written to your targets. But I can move this field over and drop it right there so it's right next to each other, so it's really clear to me, right? Make it easier to validate.
So I can use the AI prompt now that I can see my dirty data. I can use the AI prompt to go add some remediation rules and fix that if I wish. But in this case, I want to also show you and highlight this is the expression. The rule that I generated was the assert rule, which the assert rule has an expression on the left-hand side, an expression on the right-hand side, and an operator in the middle. But using our AI prompt and our AI agent that we have, I was able to generate this very complex expression with that AI prompt. And the only thing I really needed to do to be successful with that is to understand the relationships within the dataset and understand the business logic. So I'll give you another quick example of how I've used the AI prompt.
So I'm just going to put a back tick in front of this, and I'm going to say apply, which is an invalid expression. So I'll save the edit I made to my rule. It's going to come back and tell me execution failed, and so I'll tell it, "Execution failed." Probably didn't need to capitalize that. Please resolve. And so it's going to go and start debugging. It has access to our log file. It's going to go and start debugging that rule.
And so these very same concepts that I'm showing you, right? We're now looking at how we apply them to Map Designer and Process Designer, which I think is very exciting. But it's going to finish, and it's going to update the rule for me. All I have to do is say confirm. The engine's running again, and my rule has been fixed. Right? So doing that.
So I've created a couple of other examples around this kind of deeper data quality. I'm going to run through them very quickly and also just kind of skim over our data prep capabilities as well. But notice that I validated the tax rate against the state. I've got another rule in here. Let me go back to the default view in the browser. I've got another rule in here, right here, that validates revenue. So you can calculate revenue by looking at unit cost, the tax rate, discount per unit, and units sold.
Right? So I'm able to, again, use all those relationships, and now that I've theoretically, if I'd cleansed my tax rate column, I have confidence in that, right? But I can measure all of these columns using these rules. It's not just one. I can see the errors and problems in all of them. And then lastly, I use data prep to add a new field to this data set. It's called CSM, and I use the lookup to bring in and map that data into this data set and map it to this field.
And so what that's doing for me, the logic behind it is we have a customer success manager assigned to various different customers, and the criteria is that the revenue has to be at least $500,000, their subscription has to be valid, and they can't be on credit hold. So we're gluing both sides together in the sense that we're giving you the power of data quality and the standardization capabilities you need with data prep in one tool. So I've imported that data that was external to this data set into this data set, and now I'm able to run profiling and remediation against it based on that logic, which I think is just amazing. Okay, at this point, I realize I just kind of skimmed over data prep, but in a nutshell, data prep gives you the ability to make any kind of changes you want to your schema, rename your fields, delete fields, add fields. The rules that we have within data prep allow you to parse data out of a field, like address parts. If you needed to parse all the address parts out into separate fields, you could do that. You could use it to join.
You can use lookups with it. Very, very powerful. But with that, I'm going to move over to our next release, which is Integration Manager. So one of the things that we are doing within Data Connect is we are tightly coupling Integration Manager and Data Connect together. So the view that you're actually looking at is the view of our design studio. Okay? So what I'm able to do is I'm able to connect to Integration Manager and literally see all of the UIs from Integration Manager, have all the functionality and all the capabilities as if I were logged into it directly.
Okay? And so the first thing that I'm going to show you is our schedule page. So this is a reimagined view of your calendar. So all of these tiles represent different jobs. The size of the tile indicates how long the job is going to take to run. I can easily reschedule a job just by dragging and dropping. And you can see the color of the tiles is specific to the worker pool that it's associated with.
If I find a tile that's gray, that means that job has been paused. So if I click on that tile, I'm easily able to see who paused it and when, and have the ability to bring it back again. But that's just a quick view of our calendar. And next, I want to get into the observability page. So let me take a breath. When we started looking at how do we bring more value to Data Connect with our customers, specifically within Integration Manager, observability was a no-brainer. But looking at the market, what I found was there's a ton of observability tools out there, and they do a great job of visualizing, but most of them, you can't do anything with the tool, right?
It's telling you, "Hey, you got a problem," but okay, what do I do about it? How do I fix it? And all that stuff is missing. The other thing that I found is when I looked at these other tools, it's chart fatigue. It's like getting an email in all caps. You don't even know where to start and what to focus on, right? So I wanted to take a different approach to keep it very simple, and I'll explain to you what you're looking at and why we designed it this way.
So we've got these KPI cards that represent different measurements, and roughly those measurements are the health of your jobs, the health of your data, the health of the host system where our engine resides. And we've split them up into three categories: operational health, data reliability, volume and velocity. And so what you're looking at here is an aggregate. All these metrics are an aggregate of what you've got running in production. So let's say you've got 1,000 jobs running in production. You're looking at a macro view of your environment. So you can come in.
The green sparkline indicates a positive trend, the blue sparkline represents a neutral trend, and the red obviously indicates a negative trend. So you can come in here at a glance when you log in in the morning. Do I have any problems today? No, I don't. Okay, I'm going to go focus on something else. Or yes, I do. Let me drill down and figure out what's going on.
So I'll show you how it works with throughput. So throughput represents records per second in our engine. You have a filter up here that's time-based or execution-based, so you can look at the last two weeks or the last 20 executions. But effectively, all you have to do is click anywhere on this chart, and a table will appear below. So again, we're looking at records per second, so we're going to tell you what the expected records per second is for any of these different jobs, and you can see if it was successful in that or not. So again, macro to micro level, if this were a trend, you'd quickly be able to see the subset of jobs that are behind that trend. And so we take that very same concept and apply it to an actual problem.
I'll give you several examples during the rest of this demo, but next we're going to look at resource saturation. So if I click on resource saturation, I've got a historical trend chart here showing me the resource utilization on the host machine for our engine. And I've also got a visualization of queue time. So queue time means how long is the job sitting in the queue before it executes, right? So it's supposed to execute at ten o'clock directly, and how long did it sit in the queue before it actually got picked up and ran? And so I can see here I've got a negative trend with my queue time. And again, same thing, I can come and click on it, and I can see the expected queue time and the actual queue time.
Right? So this part of it will work without AI, 100%. But we're going to use these charts to help you validate what AI is telling you. AI is not always correct, and so I wanted to provide, just as we did within data quality, AI prompt is suggesting a rule to me, I have the ability to dismiss it or change it or approve it. And so I've got a trend analysis from AI telling me, "System has detected a significant increase in queue time." All I have to do is click View Metrics, and we're automatically going to filter this for you. Yes, I do have a queue time problem. I want to learn more, right?
So when you do that, we're going to have a deep scan of your metadata, which we're keeping in a database that is installed with Data Connect or with Integration Manager. I'll be able to go even deeper to understand what's going on. So key findings. Average queue time has increased to 510 seconds. These are the five jobs that are experiencing delays. And I can go, again, right back down to my chart and validate that. 510 seconds, 510 seconds, same five jobs.
Okay, what can I do about it? Immediate short-term and long-term needs. So immediate, we're recommending that you are oversubscribed in this worker pool at this time, so we're going to help you load balance, right? So very simple example of how that works. Additionally, we've got two other KPI cards over here that are used for you to manage, and resolve any kind of threshold violation that you may have or any type of failure that you may have. So a threshold violation would be a condition that you put on your job with an integration manager, i.e., I expect my job to complete in five minutes. If it takes six minutes, you're going to get a notification, either Slack, Teams, or email.
We'll direct you to this page where you can either go log a JIRA for it or assign it to somebody. In other words, you're able to manage your queue. Same thing with failures. These charts are a little bit different in the sense that we've got a legend down here that allows you to drive this visualization as well as the table below. Right over here is the native JIRA integration I was talking about. We'll have all these fields filled in for you. All you have to do is hit Create Ticket, if you're using JIRA.
But notice that I've got my mean time to resolve, which is something that Lauren referenced earlier. I can toggle between looking at an aggregate of all my failures, or I can go look at a specific failure, right? So if I'm looking at configuration errors, I'm looking at the actual in my rolling average, and note that my mean time to resolve is 13 hours. If I were to go to something else, like let's say resource issues, for example, I can see I have a different mean time to resolve. So in my opinion, I think that gives you the ability to effectively manage your queue and effectively plan quarter to quarter. How much time am I spending or are we spending, and how long is it taking to resolve any of these issues? Why is it taking us so much more time to fix this versus that?
And so that's how that works. Additionally, we've got an AI prompt up here, and so we put this in place so we don't have to clutter up our dashboard with all these different widgets and all these different options. You can literally have a conversation with your data to find out whatever you want to find out. So, for example, I'll ask it, "How many jobs were deployed last week?" And it will come back and tell me. These are the jobs that were deployed. Here's who did it, here's who owns it, here's the schedule, so on and so forth. But we put this in here specifically to give you all the flexibility that you want and that you need at your fingertips, right?
Don't wait for us to come up with a new chart that's measuring something that you're interested in. You can do it right here. Next, we're going to move over to the Configure page. So the Configure page, first thing I want to mention is we're bringing some structure in terms of how we store your configurations. So you'll be able to store all your configurations by project. So very similar to what we're doing within Repository Manager today. Additionally, you'll notice that we have the same KPI cards here.
The thing to note here is that we're context switching. When I was on the Observability page before, we're looking at all jobs. Within the Configuration page, we're looking at a single job, right? So if I'm a data steward or a data engineer, I built this solution. I may not be all that interested in everything that's going on in production, but I'm very interested in the jobs that I'm responsible for. And so KPI cards work very similarly in the sense that I can see green, blue, and red, which indicate positive, negative, and neutral. I'm going to show you how you can drill down on that in a second.
You have all your configuration details down here. You've got your list of macros, your job history. How many times does this thing run? How many times did it fail? How many times was it consecutive? We're introducing a change log, so anytime anybody touches this configuration, we're keeping record of who did it, when, and what did they do. And I'm going to show you how that factors into this in just a moment.
But next, I want to show you lineage. So let me maximize this for you. So this is a representation of lineage right here. So this is a dataset node, and this is a step node. You can click on either of them and get details. So that's showing you the schema. I can click on the step node, and that's going to give you some details about what's going on within that step.
I have the ability to search. So if I type in valid As a keyword, it will find the word valid in the data set node, it'll find it in the field, it'll find it in the step. I can also perform impact analysis. So if we had a failure here, I could click on impact analysis. It would show me the direct dependencies and show me which step the error was emanating from. Right? And so the intent of lineage here, before I get ahead of myself, you can also-- Let me make this bigger.
You can see the upstream and downstream jobs within your lineage view as well. So you can expand those. I can map fields within my project, right? So I can see how all those fields are being mapped throughout this project. But the intent around lineage for us is twofold. It's a reference for you, so if you need to do maintenance or there's a problem with this job, you can come and consult lineage first to understand how data is flowing through it, how it was designed, and what it's doing, long before you ever open up the artifact. Right?
You can also identify dependencies. So if we had a problem with this job and we needed to take it offline, we immediately understand what's upstream and what's downstream, and what do I need to do with the schedule. I think there's a ton of possibilities with this. I think auditing is going to be very useful for this. I want to have kind of a-- I don't have it visualized for you, but I want to have kind of almost a debug experience with this where you can step through and see over time how your solution has evolved. When it was deployed, it was in this state, and today it's in this state, and here's a snapshot of all the changes and who did it and what they did throughout its life cycle. So next, I want to show you, similar to what we looked at with observability, but this time we've got a KPI card here that's showing us that we've got an AI trend, and it's negative.
So let's go explore that. So here, we're going to be looking at data quality trends. And so just like before, we're getting a summary from AI saying that index declined 9% over the last week. So I can click on view metrics and our visualization, again, using our visualization to validate what AI is telling us. If this doesn't line up, I can dismiss it and go on with my day. If it does, I have the ability to do a deeper research to learn more and figure out what to do next. So it's telling me my data quality index declined.
Okay, I see it. I believe it. So let's go with the root cause analysis. And again, this is an on-demand scan to make sure that you have up-to-date information and can figure out what to do next. So it found a different pattern. The completeness dimension dropped 13%, driven by email field decline from 89% to 68%. And then it's telling me which rule failed.
So again, before I do anything, view metrics, and we're going to automatically filter all these charts. So this first chart is showing me the completeness score, and I can see the trend. This next chart is showing me the email field, which contains a rule associated with the completeness dimension. Same trend. And then finally, this rule, is not null, is configured against email. Right? Same trend.
So all three are showing me and telling me the same thing. And by the way, I now understand the root cause of my problem before I ever even opened up the artifact. And that part of it will work without AI as well. But because we have lineage and because we have that change log, I can see that my upstream job is where the problem is coming from. Field mapping missing for email field. Source field source email not mapped to target email. I can see the impact in terms of the records that were impacted, and I can also see the change history that we just looked at.
Some guy named Chris Gilson deleted the mapping. Right? So we know who to be mad at, but we also know what our problem is. And down below, you can see the job trace. So this is the job we're looking at currently. That's the upstream job, that's the downstream job. And we've got a list of suggested actions.
And so we can use this natural language prompt, just like I showed you in data quality, to make these changes. So here I'm going to say update the upstream job and restore the mapping for the email field. So when I click submit, what's happening here is we're going to show you the exact part of the UI from the upstream job where the change is being made, and you have the ability to approve it or cancel it. So in this case, I'm going to approve it. And once that's done , you'll see that that has been stricken out because we've done it already. Next, we're going to go make a change to this job, and we're going to go add a remediation rule. And this is going to fix all those null values that happened because the mapping was missing.
You can see our condition here is blank, so I'll approve it. And then lastly, we're going to put a safeguard in place, so if we ever have a problem like this again, we'll get notified. So with this, we're going to go add a notification with a completeness threshold. We'll set it at 90%, we'll approve and deploy. Right? And so this is what I was talking about, about adding some design services to Integration Manager, so every time you run into a problem, the first step isn't just downloading everything into your local repository, resolving all the dependencies, and then starting to figure out what went wrong. And I think that's very powerful.
And I'll be the first to admit, if you've got systemic problems within a massive process with 1,500 steps, this probably isn't the best avenue to go. Right? I would bring it down into the IDE and start to figure it out there. But I think the majority of the things that you guys run into, we'll be able to help you solve quickly and confidently with this. And then lastly, we have another AI prompt up here, but in this case, it's a little bit different in the sense that you can ask it questions about this job, but you can also take action. Like download and import this into my repository. Suggest better scheduling options.
A multitude of different things that you can do with this. So, that's some insight into- Our data quality release 14.0, some insight into the next release, 14.1. One thing I meant to mention at the very top of the meeting, and I didn't, and I skipped over it in my haste and my excitement, is I wanted to reassure you guys that with version 14, we have the same notion of a workspace and project as Eclipse does or did, and we are supporting that, enforcing that with version 14. So when version 14 comes out, all you're going to have to do is point version 14 to your existing workspace. There's no upgrade, there's no migration, just point it to your existing workspace and away you go. Okay, that does it for me. Thank you.
Thanks, Chris. I know we went over a whole lot. I hope everybody listening sees and is excited as we are about the things that are going to the product. And hopefully you guys can see how we're trying to approach it from both the technical lens, as well as, again, that business value driver, right, where we can make it easier to use and operationally efficient. So with that, we're going to go to the Q&A and address a few questions that came up. The first one that's in here is about remediation changes. So the question is, "Do these remediation changes get pushed to the source system?
And if not, where is the remediation data stored?" That was answered here, that remediation builds the rules to make those changes, and then you have the option to direct that remediation to change it at the source or to intervene on that. Yeah. So remediation, and that's a good question. So we're never changing the source. We never change the source in any of our tools, and we will never change the source in the future. The remediation happens when you write it out to a target, right? So you're reading it from something, and you want to write it to something.
And you have the ability within Data Connect to suppress the targets if you wish, if you just want it to generate metrics. You also have the ability to suppress the remediation if you don't want it to be remediated. So hopefully that answers that question. Yeah. Perfect clarification there. Another question we received, which you just hit on, Chris, is the upgrade path, and so I don't think I necessarily need to reiterate that. I will say that, as I mentioned at the beginning when we kicked off the webinar, we're really trying to continue to lean in, right?
We want our partnerships with all of you to continue and to be strong, and part of that is just making sure we're in regular engagement and doing regular assessments on how you've deployed, how you're using the tool, what is and isn't working. And so as you start to evaluate V14, and hopefully, again, are as excited as we are to utilize it, don't be shy in making sure that we come around you to make sure that's a smooth process for you. Again, I think we've designed it so that it will be, but there's also all of the other business drivers that we can help you guys accomplish as well. Anything you'd add to that, Chris, or anything that you want to highlight outside of that? Yeah. We're hyper-aware of the fact that the most important thing to our users is all the collateral that they've built, all the years that they've spent building solutions and maintaining them, and we certainly want you to continue to extract value from those, and we don't want to make that process painful for you in any way, shape, or form. We just want to be very accommodating to that.
Let's see. Another question here. Let's see. Going through this. So there's a question here on the availability of V14. So this is coming-- Obviously we're here debuting it to you guys, so very short order. Looking at end of the month, we are going to make more of a public launch, or relaunch, if you will.
So please go to our website for that. Some of the content is already on there. It will continue to get updated. And again, we're wanting to make sure that we're leaning into this. So would love any feedback as well. We want to make sure that what we're excited about, you're equally excited about. So, we also have a number of programs internally.
We are starting a central chapter. We'll hopefully have an east and a west as well. Those are all meant to be forms of feedback, right? So again, as we continue to iterate on V14, we would love for you guys to make sure you're looking at the release, giving the feedback where you can, reaching out to your rep in your area. If you don't know who that is, don't hesitate to reach out to myself or even Chris. I know either one of us are more than happy to assist you there. And then I think that is it.
Let me just double-check, make sure there aren't any other questions in the chat here. One second. Yeah. So I think that wraps on the questions, unless anybody else wanted to put anything in here last minute, you're more than welcome to do so. Lauren and Chris, we do have a couple that came in. Are there any changes in deployment options in V14? Yes, and that comes with the inclusion of Integration Manager.
So Integration Manager, you can deploy it anywhere. It can be deployed on-prem, it can be deployed as a container in a VPC. You can deploy it to the cloud. You can put our engines anywhere. But in terms of our design studio, we wanted to keep the design studio private, desktop, single tenant. Excuse me, Integration Manager is multi-tenant. We do have the ability, from an architecture perspective, to deploy the design studio to the cloud, but right now, all we're focused on is on-prem.
Okay. And there was one more. A lot of organizations have in-house AI systems, so when you mentioned that you could use AI in there, there were some questions around what's supported there, so on and so forth. Yeah, great question. Yes, you can plug your own model into V14. Absolutely. Excellent.
And with that, I don't see any other questions that have come in. All right. Then I guess with that, we can certainly wrap. Again, I want to make sure just to say thank you to everybody that took the time to attend this. We've barely scratched the surface, I feel like. There's a lot of really exciting things in what we're doing in V14. It's the direction we're going within the product itself, and everything always comes back to how trusted is your data, and this is really enabling us to be proactive, reactive, do it in a single lens, if you will, do it more efficiently and more effectively.
So hopefully you guys see that and take that away, and we look forward to connecting with you live. Appreciate it. Thanks, guys. Thanks.