Case Study 3: Professional Accountant in a Large Firm – Synthetic Client Data for LLM TrainingCase Study 3: Professional Accountant in a Large Firm – Synthetic Client Data for LLM TrainingCase Study 3: Professional Accountant in a Large Firm – Synthetic Client Data for LLM TrainingCase Study 3: Professional Accountant in a Large Firm – Synthetic Client Data for LLM Training
  • Home
  • About
    • Education and Training
    • Board of Directors
  • Values
    • Tackling Economic Crime
    • Contribution of the Profession
    • Regulation
    • Ethics Group
    • SORP LLPs
  • Ethics & Insights
    • CCAB Ethical Leadership Podcast
    • Ethics Resources
    • Ethical Dilemmas Case Studies
    • AI Hub
  • Podcasts
  • News
  • Contact
✕

Case Study 3: Professional Accountant in a Large Firm – Synthetic Client Data for LLM Training

7 May, 2026

SCENARIO

You are a partner responsible for innovation and technology at a Big Four accounting firm. The firm is developing its own large language model (LLM) to enhance various service lines, including audit, tax, and advisory. The technology team has proposed creating a synthetic dataset derived from actual client data to train the LLM, arguing this approach would significantly improve the model’s performance while addressing direct confidentiality concerns.

The proposed process would work as follows:

  • Client financial data, transaction records, contracts, and other sensitive information would be collected from across the firm’s client base.
  • This data would be used to generate “synthetic” data that maintains the statistical properties and patterns of the original data but doesn’t contain actual client information.
  • The synthetic data would then be used to train the LLM, which could subsequently assist professionals with tasks like contract analysis, anomaly detection, and generating first drafts of reports.

The technology team argues that since the actual client data isn’t directly used to train the LLM, but only to create synthetic data, client confidentiality is maintained. They’ve also noted that competitors are developing similar capabilities, and the firm risks falling behind if it doesn’t move forward with this project.

However, several concerns have been raised:

  • While the synthetic data wouldn’t contain direct client information, it might still retain patterns unique to specific clients that could potentially be extracted or inferred by sophisticated users. Are there sufficient client numbers and industry spread to make data sets representative? This could also result in data revealing information about clients.
  • The firm’s engagement letters and terms of business don’t explicitly mention using client data, even in anonymised or synthetic form, for training AI systems.
  • Some team members have questioned whether the distinction between using synthetic data versus actual client data is meaningful enough from an ethical and legal perspective.
  • There is uncertainty about whether clients from regulated industries (like financial services or healthcare) would have specific restrictions that might prohibit even this indirect use of their data.
  • The technology team has acknowledged that synthetic data generation isn’t perfect, and there’s a small but non-zero risk that actual client data could occasionally “leak through” into the synthetic dataset.
  • Some partners have raised concerns about whether accepting client authorisation to use their data might violate the FRC Ethical Standard’s provisions on gifts and favours (paragraph 4.38).

FRC and IESBA Guidance Recent guidance from the Financial Reporting Council (FRC) has clarified several important points relevant to this situation:

  • The IESBA Code requires authorisation from the audited entity before any confidential information is used in the development of technology.
  • Even when client data is processed to create derived or synthetic data, authorisation from the client is still required. This is because the first step in that process (processing the original client data) constitutes using confidential information for the development of technology, albeit as a preliminary step.
  • The FRC has clarified that accepting client authorisation to use their data for technology development is not prohibited under the “gifts and favours” provisions of the FRC Ethical Standard (paragraph 4.38).
  • The FRC suggests that clients might be more amenable to authorising the use of their information for creating synthetic data (rather than direct use) as long as their confidentiality is protected.
  • Creating synthetic data through the application of patterns observed by human expertise over time (without directly processing client data) would be permitted under the IESBA provisions.

Ethical Considerations

Integrity

  • Would proceeding with processing client data to create synthetic data, without explicit client authorisation as required by the IESBA Code (and highlighted by FRC guidance), be consistent with being straightforward and honest in professional relationships?
  • Does relying on the technical distinction between synthetic and actual data, despite the residual risks of leakage or pattern inference, fully uphold the spirit of transparency and integrity owed to clients regarding the use of their information?

Objectivity

  • Could competitive pressure to match rivals’ AI capabilities be unduly influencing the firm’s assessment of the ethical risks and the importance of client consent, thereby compromising objectivity?
  • Are you objectively balancing the potential commercial benefits of the LLM against the ethical obligations, particularly concerning client confidentiality and data usage rights?

Professional Competence and Due Care

  • Does the firm fully understand the technical limitations and risks associated with synthetic data generation (e.g., potential for data leakage, re-identification)? Have reasonable steps and due care been taken to mitigate these risks?
  • Are you demonstrating professional competence by staying abreast of, and adhering to, the latest regulatory and ethical guidance (IESBA Code, FRC clarifications) regarding the use of client data for technology development?

Confidentiality

  • Does the act of processing original client data to create synthetic data constitute a use of confidential information requiring explicit authorisation under the IESBA Code, even if the final synthetic dataset contains no direct client identifiers?
  • What are the firm’s obligations to protect client confidentiality when using their data for purposes (LLM training) beyond the scope of the original engagement, and how can these be met robustly?

Professional Behaviour

  • Could proceeding without explicit client consent, even for synthetic data creation, damage the firm’s reputation and public trust in the profession if it becomes known, potentially violating the spirit of ethical conduct?
  • Does the proposed approach fully comply with both the letter and spirit of data protection laws (e.g., GDPR) and professional standards regarding data usage and consent?

Possible Course of Action

  • Seek explicit client authorisation: In line with IESBA Code requirements 114.2b and 114.3 A3, obtain explicit authorisation from clients before using their data, even for the purpose of creating synthetic data. The FRC has clarified that accepting such authorisation does not violate the “gifts and favours” provisions.
  • Conduct a comprehensive risk assessment: Engage independent technical and legal experts to evaluate the synthetic data approach, identifying potential risks of data leakage or re-identification.
  • Update client agreements: Revise engagement letters and terms of business to explicitly address the use of anonymised or synthetic data derived from client information for AI training purposes.
  • Implement a client consent framework: Develop a tiered approach where clients can opt in or out of having their data used for synthetic data generation, with clear explanations of the process and safeguards.
  • Establish data governance protocols: Create strict governance procedures for the selection, transformation, and use of client data, including appropriate anonymisation techniques and security measures.
  • Develop robust testing procedures: Implement rigorous testing of synthetic datasets to ensure they don’t contain actual client information or allow for re-identification of clients.
  • Create an ethical oversight committee: Establish a committee including partners from different service lines, ethics specialists, and potentially external advisors to review and approve AI data usage practices.
  • Implement a phased approach: Begin with a limited pilot using only data from clients who have given explicit consent, allowing for evaluation and refinement of safeguards before wider implementation.
  • Consider alternative approaches: Explore whether synthetic data could be created based on patterns observed through human expertise rather than directly processing client data, which may be permitted without specific client authorisation according to the FRC guidance.

Recommendation

Strategic approach:

  • Recognise that under both the IESBA Code and recent FRC guidance, client authorisation is required before using their data to create synthetic data, regardless of how transformed the final data may be.
  • Proceed with caution, acknowledging that the distinction between synthetic and actual data may not be meaningful from an ethical or regulatory perspective without appropriate safeguards and client authorisation.
  • Develop a clear communication strategy for clients that explains in plain language how their data would be used, the safeguards in place, and the benefits to service quality. As the FRC suggests, clients may be more receptive to authorising the creation of synthetic data than direct use of their information.
  • Implement a formal opt-in process for existing clients rather than assuming implied consent, and update standard terms for new clients to explicitly address potential uses of their data for technology development.
  • For clients in highly regulated industries, conduct specific legal reviews before including their data in the synthetic data generation process, as they may have additional restrictions beyond the general ethical requirements.
  • Establish an ongoing monitoring programme to regularly test synthetic data for potential client information leakage, with clear protocols for addressing any issues that arise.
  • Document these decisions and processes thoroughly to demonstrate due care and professional judgement in alignment with both the letter and spirit of the ethical standards.
  • Consider the alternative approach of creating synthetic data based on patterns observed through human expertise rather than directly processing client information, which may provide a path forward that doesn’t require specific client authorisation.
Share
Subscribe to our Newsletter
Privacy Notice
Links
How to choose an Accountant or Tax Advisor
Contact Us
© 2020 CCAB. All Rights Reserved | Privacy Policy
Registered in England and Wales No. 1864508
Registered Address: CCAB Limited, Chartered Accountants' Hall, Moorgate Place, London EC2R 6EA