TY - GEN
T1 - Slim-Donut
T2 - 2025 Supercomputing India, SCI 2025
AU - Dhondge, Om
AU - Bayyapu, Neelima
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - The large model size and high computational cost of the Document Understanding Transformer (Donut) and other OCR-free architectures from the Visual Document Understanding (VDU) models pose significant obstacles to scalable and practical deployment. In this paper, we present Slim-Donut, a new lightweight and scalable framework designed for high-throughput generation of ToCs (Table of Contents). Slim-Donut optimizes the standard Donut architecture through a two-pronged approach: first, it incorporates a dynamic token slimming module into the Swin Transformer encoder to reduce computational complexity by adaptively merging redundant visual features; second, it employs post-training quantization and knowledge distillation to significantly compress the model size while preserving high fidelity. By fine-tuning the comprehensive DocLayNet dataset, Slim-Donut learns to accurately identify and structure hierarchical headers. Our experimental evaluation shows that Slim-Donut achieves an 8 × improvement in CPU inference speed, a 3 × reduction in model parameters, and a 2.5 × reduction in GFLOPs compared to the baseline Donut model, while maintaining comparable accuracy. Furthermore, it outperforms other leading architectures like LayoutLMv3 in scalability and deployment readiness. Crucially, this efficiency is achieved with only a minimal trade-off in extraction accuracy, establishing a new, highly favorable point on the performance efficiency frontier for document indexing tasks. This work demonstrates the immense potential of architectural optimization and model compression to create practical, scalable solutions for automated document structuring.
AB - The large model size and high computational cost of the Document Understanding Transformer (Donut) and other OCR-free architectures from the Visual Document Understanding (VDU) models pose significant obstacles to scalable and practical deployment. In this paper, we present Slim-Donut, a new lightweight and scalable framework designed for high-throughput generation of ToCs (Table of Contents). Slim-Donut optimizes the standard Donut architecture through a two-pronged approach: first, it incorporates a dynamic token slimming module into the Swin Transformer encoder to reduce computational complexity by adaptively merging redundant visual features; second, it employs post-training quantization and knowledge distillation to significantly compress the model size while preserving high fidelity. By fine-tuning the comprehensive DocLayNet dataset, Slim-Donut learns to accurately identify and structure hierarchical headers. Our experimental evaluation shows that Slim-Donut achieves an 8 × improvement in CPU inference speed, a 3 × reduction in model parameters, and a 2.5 × reduction in GFLOPs compared to the baseline Donut model, while maintaining comparable accuracy. Furthermore, it outperforms other leading architectures like LayoutLMv3 in scalability and deployment readiness. Crucially, this efficiency is achieved with only a minimal trade-off in extraction accuracy, establishing a new, highly favorable point on the performance efficiency frontier for document indexing tasks. This work demonstrates the immense potential of architectural optimization and model compression to create practical, scalable solutions for automated document structuring.
UR - https://www.scopus.com/pages/publications/105035728779
UR - https://www.scopus.com/pages/publications/105035728779#tab=citedBy
U2 - 10.1109/SCI68648.2025.11333857
DO - 10.1109/SCI68648.2025.11333857
M3 - Conference contribution
AN - SCOPUS:105035728779
T3 - 2025 Supercomputing India, SCI 2025
BT - 2025 Supercomputing India, SCI 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 9 December 2025 through 13 December 2025
ER -