Template-Grounded Generative AI for Official Correspondence: Design and Functional Evaluation of the Smart Correspondence Assistant
DOI:
https://doi.org/10.37385/ceej.v7i3.11864Keywords:
Design Science Research, Generative Artificial Intelligence, Official Correspondence, Template-Grounded Generation, Functional EvaluationAbstract
Official correspondence in professional associations often relies on copied files and separate approval records, producing inconsistent structure and weak traceability. General-purpose generative artificial intelligence can accelerate drafting but may fabricate factual content and vary mandatory document elements. This study designs and functionally evaluates the Smart Correspondence Assistant (SCA), a web-based artifact that uses PermenPANRB 21/2021 as a structural reference while separating deterministic document assembly from model-generated narrative. Following design science research, seven design requirements were derived from practitioner experience, organisational letter examples, and the reference standard. The artifact supports seven letter types, configurable multi-level approval, deterministic numbering, role-based access control, audit logging, and PDF generation. Functional evaluation comprised fifty-four scenarios across six iterative runs in an isolated PHP 7.4.33 environment. Eight defects were identified and all were remediated and retested. In the final cumulative state, all fifty-four scenarios passed: forty-six were decided through objective functional checks and eight semantic or visual scenarios were verified manually by the first author using preserved evidence packages. Because the manual reviewer was part of the development team, the result represents developer-executed functional conformance rather than independent validation. The study contributes a transferable pattern for constraining generative artificial intelligence in official correspondence by keeping structural and authority controls at the application layer, while treating generated narrative as reviewable content rather than a final document.
References
[1] O. R. Danar, "Digital transformation of Indonesian administration and bureaucratic system," International Journal of Electronic Governance, vol. 16, no. 2, pp. 152-171, 2024, doi: 10.1504/IJEG.2024.140789.
[2] B. Kusumasari, "Digital democracy and public administration reform in Indonesia," International Journal of Electronic Governance, vol. 10, no. 3, pp. 317-337, 2018, doi: 10.1504/IJEG.2018.095937.
[3] Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, "Survey of hallucination in natural language generation," ACM Computing Surveys, vol. 55, no. 12, art. 248, pp. 1-38, 2023, doi: 10.1145/3571730.
[4] S. Nzobonimpa, J.-F. Savard, and J. Lawarée, "Generative AI in public administration: evaluating a fine-tuned large language model for policy briefing notes," Science and Public Policy, 2026, doi: 10.1093/scipol/scag026.
[5] J. Maynez, S. Narayan, B. Bohnet, and R. McDonald, "On faithfulness and factuality in abstractive summarization," in Proc. 58th Annu. Meeting Assoc. Comput. Linguistics, 2020, pp. 1906-1919, doi: 10.18653/v1/2020.acl-main.173.
[6] N. Haug, S. Dan, and I. Mergel, "Digitally-induced change in the public sector: a systematic review and research agenda," Public Management Review, vol. 26, no. 7, pp. 1963-1987, 2024, doi: 10.1080/14719037.2023.2234917.
[7] I. Mergel, N. Edelmann, and N. Haug, "Defining digital transformation: results from expert interviews," Government Information Quarterly, vol. 36, no. 4, art. 101385, 2019, doi: 10.1016/j.giq.2019.06.002.
[8] A. R. Hevner, S. T. March, J. Park, and S. Ram, "Design science in information systems research," MIS Quarterly, vol. 28, no. 1, pp. 75-105, 2004, doi: 10.2307/25148625.
[9] K. Peffers, T. Tuunanen, M. A. Rothenberger, and S. Chatterjee, "A design science research methodology for information systems research," Journal of Management Information Systems, vol. 24, no. 3, pp. 45-77, 2007, doi: 10.2753/MIS0742-1222240302.
[10] S. Gregor and A. R. Hevner, "Positioning and presenting design science research for maximum impact," MIS Quarterly, vol. 37, no. 2, pp. 337-355, 2013, doi: 10.25300/MISQ/2013/37.2.01.
[11] J. Laux, "Institutionalised distrust and human oversight of artificial intelligence: towards a democratic design of AI governance under the European Union AI Act," AI & Society, vol. 39, no. 6, pp. 2853-2866, 2024, doi: 10.1007/s00146-023-01777-z.
[12] B. Green, "The flaws of policies requiring human oversight of government algorithms," Computer Law & Security Review, vol. 45, art. 105681, 2022, doi: 10.1016/j.clsr.2022.105681.
[13] H. Zhang, H. Song, S. Li, M. Zhou, and D. Song, "A survey of controllable text generation using transformer-based pre-trained language models," ACM Computing Surveys, vol. 56, no. 3, art. 64, pp. 1-37, 2024, doi: 10.1145/3617680.
[14] Y. Jiang, Y. Wang, X. Zeng, W. Zhong, L. Li, F. Mi, L. Shang, X. Jiang, Q. Liu, and W. Wang, "FollowBench: a multi-level fine-grained constraints following benchmark for large language models," in Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics (Vol. 1: Long Papers), 2024, pp. 4667-4688, doi: 10.18653/v1/2024.acl-long.257.
[15] Z. Li, B. Peng, P. He, and X. Yan, "Evaluating the instruction-following robustness of large language models to prompt injection," in Proc. 2024 Conf. Empirical Methods Natural Language Processing, 2024, pp. 557-568, doi: 10.18653/v1/2024.emnlp-main.33.
[16] A. Hua, K. Tang, C. Gu, J. Gu, E. Wong, and Y. Qin, "Flaw or artifact? Rethinking prompt sensitivity in evaluating LLMs," in Proc. 2025 Conf. Empirical Methods Natural Language Processing, 2025, pp. 19889-19899, doi: 10.18653/v1/2025.emnlp-main.1006.
[17] Ministry of Administrative and Bureaucratic Reform of the Republic of Indonesia, "Regulation No. 21 of 2021 on Official Correspondence within the Ministry of Administrative and Bureaucratic Reform," 2021. Available: https://peraturan.bpk.go.id/Details/170616/permen-pan-rb-no-21-tahun-2021
Template


