Overview

A mid-sized consulting firm was grappling with a growing document-management problem. Over the last few years, departments have amassed about 10,000 files in different formats, such as Word documents, Excel spreadsheets, CSV spreadsheets, images, text documents, HTML pages, JSON records and XML documents. The files contained client projects, financial reports, internal documents, presentations and archived records.

The management team wanted to move these records to PDF because it was easier to share, archive, and review without having to worry about the original application that was used to edit them. But converting such a large and mixed collection was a challenge.

The team tried manual methods at first, then looked at web-based converters. Neither option provided the efficiency, privacy and consistency they wanted. Finally, they tested a desktop-based solution and developed a repeatable workflow for the whole archive.

What Is Mixed File to PDF Conversion?

Conversion of files of different formats into PDF format is known as Mixed-file-to-PDF conversion. Instead of creating separate files in Word, Excel, image, HTML, JSON, XML, and text formats, an organization can always generate PDF copies for standard viewing and archiving.

For this project this was important because it was common for the employees to have to look at older documents without knowing what application created it in the first place. A standardized PDF format also helped to organize the project records into central folders.

The application supports over 20 formats including DOC, DOCX, TXT, RTF, Markdown, JPEG, PNG, BMP, TIFF, GIF, HEIC, SVG, XLS, XLSX, CSV, HTML, HTM, MHT, JSON and XML.

Initial Approach: Manual File Conversion

The company administration team initially attempted to convert the files one by one.

An employee would then open the document, select print or export, select PDF, set up the output, save the file, and move it to the appropriate archive folder. We applied the same procedure for spreadsheets and images.

At first it seemed doable. But after a few hundred files, the workload became hard to handle.

The team had some problems: