Skip to main content
AI Safety Forum Australia
talkIndustry, adoption & educationCross-cutting

Responsible AI Adoption in Public Sector Auditing

8 July 2026 · 3:30 pm–3:55 pm · Refectory

Generative AI has significant potential to enhance public sector service delivery, but its adoption must be supported by robust governance and responsible AI practices. This presentation shares practical lessons from an ongoing collaboration between CSIRO and the Audit Office of New South Wales (AONSW) on applying generative AI in real-world auditing workflows. The talk explores how responsible AI principles can be translated into practice, including human oversight, trustworthy evidence retrieval, governance, and risk management.

Recording

Speaker

Audience Q&A

Ask a question or upvote others.

Loading questions…

Transcript

0:03

Hi everyone. My name is Ming Ding. I'm currently looking after the Privacy Technology Group at the CSIRO. I also hold adjunct professor positions at Swinburne University of Technology, University of Technology, Sydney. I'm very pleased to be here and would like to share our current experience with a few government projects in terms of AI adoption. In particular, there is an ongoing collaboration between CSIRO and Audit Office for New South Wales. I just want to share the lessons learned of the experience, the practical AI tools that we have built.

0:49

Hope that will be useful. First, a very quick connection about my talk to the international AI safety report. The idea of our work is clearly categorised as in risk mitigation and technical safeguards. And the monitoring here is the close. That is highly relevant. So we develop solutions that can monitor, control and adjust AI model behaviours in real-life deployment. So here is a quick outline of my talk today. First I will give a very quick introduction about my group. And then I will provide more context about adoption in the government sector.

1:31

And the last part, it would be the focus of my talk. I will share the tools developed together with our Audit Office for New South Wales. So very quickly, our group is the Privacy Technology Group at the CSIRO. Our vision is very simple. We think data and AI systems, they should go hand in hand, and they should be designed in a holistic way so that they are safe, privacy preserving and secure. So as a result, we have two teams. One is the data privacy team, taking care of the bottom layer.

2:10

For the moment, we are working with Home Affairs on developing a data classification framework so that organisations can develop a common language to understand the risk value and the guardrail methods for sharing data. Second team on top. That is the private and confidential AI team, focuses on building trustworthy systems, including privacy-aware, fairness-aware AI models and distributed learning and machine unlearning as well. So we will not go to details today. So here is a link to our website and the QR code. Feel free to access our web page or to find out more about us.

2:55

Okay, so next I would like to give a little bit more context about AI adoption in the public sector. First I would like to mention two national initiatives last year regarding cross-sector AI adoption. First one is Australian public sector AI plan, released last November. So there are three pillars: trust, people and tools. So trust means AI ethics and governance, currently led by DTA, Digital Transformation Agency. Second pillar, people, basically meaning staff capability building and training. So we are currently working with our Australian Public Service Commission on delivering a pilot AI training program across the federal government agencies.

3:41

The third pillar is the tools. So basically means tooling and applications across the government sector. So this pillar is currently led by the AI team at the Department of Finance. So I'm currently serving as a member of the working group at GovAI. So we are looking at how to enable AI application tool development and sharing across as a government to sectors. And there is a second initiative which is called Australian National AI Plan, targeting the private sector. So there are also three pillars: capture the AI opportunities, spread the productivity gains or benefits, and finally keep the adoption and implementation of AI safe and responsible.

4:33

So here are a few real-life incidents that remind us why AI governance and safety matter in reality. So first, last year there was a company called Ross Intelligence. That company lost a landmark lawsuit to Thomson Reuters because they have used the copyrighted material for training their AI model. As a result, Ross Intelligence has ceased operations due to high legal cost. Second, last year there is a New South Wales subcontractor uploading people's sensitive data onto ChatGPT, including thousands of people's financial data, house information, etc. And this particular incident has been identified and investigated as a data breach by New South Wales Cyber Security and the

5:26

New South Wales Reconstruction Authority. Finally, last October, a consultation company had to refund a government agency, because in their government, they found in the report there are many in the hallucinated reference that misquote. So these real-life examples, it shows the AI safety can be directly linked to legal risk, data protection and hallucination, those kind of real-life risks. The next slide shows the ethics principles. In 2019, eight AI ethics principles were established by the Department of Industry. So for today's talk, there is no need to go through the details, because the principles include many aspects including privacy, security, fairness, traceability, transparency, accountability, you name it.

6:27

So in reality, the most important question is how to realise these principles through policies and technical solutions. So regarding policy side. So the first thing I would like to mention is the recent pieces of documents or guidelines released by the Department of Industry. We have been working with the Department of Industry for more than one year. So the first piece is a guidance for AI adoption. And the second one focuses on transparency and the proper disclosure of AI-generated content. The Department of Industry's guidance, mainly targeting the private sector, but the

7:10

for the government sector, this is the policy released by DTA last December. This is called a policy for the responsible use of AI in the government. So after the release of the policy to also release that the timeline of AI adoption in our government. For example, last month, the federal government agencies, they all released their positions of AI adoption on their website. And by the end of this month, every agency needs to appoint a chief AI officer. And by the end of this year, each federal-level government agency needed to create an AI use case and register, establish a responsible AI approach, and implement staff training and, more importantly, start working on an AI use case assessment.

8:03

So the expectation will be all of those paperwork and assessment and approval needed to be done by April next year. So that's the timeline and our policies set out by DTA. So the next part, I will use our real-life collaboration with the Audit Office for New South Wales to showcase how we work with government agencies to implement the policy and develop real-life AI solutions. So first, this collaboration with the Audit Office is a multi-year partnership. The basic idea is to create a platform so that we can share knowledge.

8:43

For example, from the scientist part, we share the latest in the development of AI. From the Audit Office, they share their workflows of their daily work. And we co-design AI applications. And the fundamental goal is not to replace auditors but to enhance their productivity and our collaboration. That has been reported by a few news outlets listed here. And here I show a picture of the Auditor-General for New South Wales, Bola Oyetunji, and our research director, Doctor Liming Zhu. So here I just want to mention that this particular collaboration actually started with a vision, a dream from Bola, the Audit Office for New South Wales.

9:29

So he was interested in exploring something called predictive auditing. Nowadays, when the audit office audits the local councils, usually that's when the problems have already occurred. And the Bola would like to check whether it is possible to explore some early warning signs, some expenditure patterns of the local council so that he can approach the local council at some day later saying, hey, please be careful, because your belly is going to turn up in two years, because of this or that data, something like that. So that's the vision.

10:07

So in order to realise this vision, we developed a framework of AI use cases. We deliberately keep it general, which can accommodate many, many possibilities. So there are three categories. The first one is called 1 to 1 transformation. We start with one document or report, and we apply a set of rules, criteria. And to the end we are going to get another report. So this is very useful for use cases such as document review, policy conformance checking, automatic redaction of sensitive information. The next category is a little bit more complicated, called many to one transformation.

10:50

So we started with multiple documents. And we would like to end up with a consolidated report. So this is useful to generate knowledge or policy synthesis, literature review or evidence, fact extraction. Finally, we have the most complicated one, which is called many to many transformation. So in this case we are not dealing with the reports. We are dealing with items, artefacts. So the goal is to prioritise or rank or give scores to those items to facilitate decision-making. So this use case is the most complicated one, but it can enable very useful use cases such as content monitoring, evidence reranking, risk mitigation, prioritisation, those kind of things.

11:41

So the key point here is that this framework is not just an academic classification one. It's actually a device. The leadership level about the complexity and the governance of your use cases. From category one to category three, you can see that the complexity goes higher. And also it has the potential to bring higher value to your organisation. But of course you need to keep an eye on the more complex the governance and the risk mitigation control mechanisms. So here I would like to answer a common question you might have.

12:19

So regarding the mentioned that AI use cases, you might have wondering why not just using chat bot, such as Claude, ChatGPT, Gemini, to do that. So the answer is yes. Technically speaking, you can do that. But there are limitations. So let me use a tangible example to showcase those points. So the example is a document review. Very simple. You have one document and you would like to get a review assessment. You can use a chat bot. Typically, as a user, you upload a document and a prompt, possibly with a set of assessment rules.

12:56

So that's it. AI will send you the review report. But if you look at this setting, you will realise there are a few things missing. First, are you allowed to upload a document to the cloud? That's just to violate the privacy contractual agreement. Second, if you are using one AI, does that introduce some kind of bias? Because AI — Gemini, ChatGPT — they are all trained on different methods. So we need to keep that in mind. So there are many other limitations as well. So at the end we need to develop an AI application which considers governance adaptation, multi AI collaboration and adding human experts knowledge

13:40

or inputs into this application as well. So here I'm just showing you our homework. And at the end of the day we rounded up seven use cases. So, no time for discussing details. But just want to mention we assess those seven use cases from six dimensions, including value to auditors. Technical feasibility means whether it is easy and cost effective to implement the solution. Operational readiness means whether auditors are happy and ready to use. Security, safety, privacy. There are many goals as well. Capability upscaling means that we are not targeting use case useful for individuals, but on an organisational level.

14:25

Finally, there was a use case that needed to be extensible and scalable to other IT use cases as well. So at the end of the day, we selected the advanced predictive modelling as the final use case to go ahead, because that use case is also aligning with our original vision. A predictive for auditing. A few quick notes. First, project cost is irrelevant at a later stage, because cost depends on the scope of your MVP. So you can have a 50 K project all the way to a 500 K or even millions of dollars project, all based on the same product.

15:04

It all depends on what key features are you going to make. Second thing, we prioritise use cases that score very high across the board. Finally, two real-life projects have been completed last two months with Audit Office. So I will showcase a typical tool in the following slides. So this tool is called GUIDE — Guided Understanding and Insight for Document Exploration. So this tool is based on our previous work with Home Affairs on a six-year project and based on our submitted AI use cases to the Department

15:42

of Finance. So the title is AI for document assessment. So the idea is very simple. You are dealing with hundreds of documents. You don't have time to go through them all. So you use AI to extract those key information and evidence for your questions. And to the end you ask AI to generate a synthesis report. So that's the typical setting up of these tools. So some of you might have already got the idea. So this is the typical setting for retrieval-augmented generation, shortened as RAG. And here I show the workflow of our tool.

16:18

So first, as a user you submit your prompt together with hundreds of PDF files as a RAG system, and then you start interrogating the system and provide your input, etc., and the system, sometimes it will rewrite your prompt, under the hood. And after that, the AI system will retrieve relevant information using embedding-based hybrid retrieval. After that we have some human oversight step, which will be highlighted later. And follow the with that step, we have a large language model generating answer and some mechanisms to double check the quality of the answer, guardrails, etc. After you get to the answer, you are allowed to continue the conversation

17:04

with the AI. And is it so that you can refine your answers only. And so that's a very quick overview of the workflow. I just want to mention that. So this workflow is quite comprehensive and complicated, but actually, technically speaking, it's simple. Regarding steps three and six, this is the technical sheet you need to look at. There are ten technical decisions that you need to make. So usually if you have a team working on hundreds of hours, you can do a quick optimisation and come up with a very good design.

17:40

But the key point I would like to make today is that even with such a complicated system, it does not meet the fundamental requirements of Audit Office, because they have a few additional strategic priorities. So first one is to ensure the auditors can come in and review the AI outputs before the final output. And this is called meaningful auditor accountability and human oversight. Second one, the goal is not to replace the auditors. The goal is to augment their work. Finally, which is also very interesting, Audit Office needed to train the next generation of auditors.

18:20

They are not looking for automation, because that's useless for the next generation of junior auditors, they would like to use this tool for their day-to-day, I think, the learning and the training. So here is the solution. So that's why we have a step number four, human oversight. So let me showcase how this step number four works. So very quickly, a few screenshots. So these are standard one. So first, as a user you submit your prompt, upload your files, and then this is the step of where human oversight matters.

18:56

So we pause. When AI extracted the evidence we don't allow the AI to go ahead, this generate answer. Instead we ask the auditors that comes in and reviews there a 50 or 30 evidence AI retrieved, and the auditors can review those, the evidence, checking, verifying whether they are good and accurate, and then de-select or reorganise. You can drag as it was evidence here. And as they are saying, number seven is the most important one. And I don't want to number 11 and 13, you can do that.

19:30

So after that you allow AI to generate the answer. And of course we have a multi AI to give the answer. And you can compare and add to your inputs at it. And you'll get to the answer and we log everything. So this is another requirement from audit office. So they need a clean, transparent audit trail Due to time limitation, just a few quick concluding remarks. So AI adoption in government is basically an organisational readiness and capability-building challenge. So AI model and the tool are not the most important part.

20:02

So how to make sure the AI solution add values to the organisation is the most critical question. Second, leadership is critical. In our case, thanks to Bola's vision and his support, we have the privilege of working with Audit Office and developing those tools. So finally, human oversight is critical, because it facilitates the training of the next generation of the staff and helps us secure people's real-life job in audit office. So I will stop here and happy to take any questions. Thank you.