AI Resume Parser: Seven Formats, One CRM Workflow
An AI resume parser moved applicant data from inbound email into an existing Rails CRM without adding a separate review screen. It handled seven document formats, malformed PDFs, scanned files, and source-specific links inside one production route. How do you automate that workflow without bypassing the rules that already make a CRM contact valid?

| Project detail | Value |
| Client | An US insurance marketing platform |
| Industry | Insurance |
| Status | Live in production; no feature flag |
| Initial delivery | Roughly three weeks to production hardening |
| Refinement | Continued in later months |
| Core stack | Ruby 3.4, Rails 8.1, Sidekiq 8, OpenAI Responses API, PostgreSQL, AWS |

— The initial production route covered seven resume formats in roughly three weeks; later work refined the same workflow.
01. Summary
The AI resume parser converts an application email into a structured contact inside the client’s CRM
Mailgun receives the message, Rails determines the business unit and owner, Sidekiq handles the slow work, format-specific tools extract the resume, and OpenAI Structured Outputs returns data that matches the CRM schema. A parser demo usually starts with a clean PDF and ends with JSON. This project started earlier and ended later. Applications arrived as email notifications from Indeed and LinkedIn. The resume could sit behind an expiring link, use a legacy document format, contain no usable text layer, or fail ordinary PDF extraction. The output then had to pass the same validations and permissions as a contact created by a person. The model call wasn’t the product. The production work sat around it: source detection, ownership, document recovery, schema enforcement, CRM domain logic, and failure handling. Successful jobs run unattended. Failed jobs are logged for manual follow-up because replaying the whole job could create a duplicate contact.
02. Problem
The recruiting business units needed applicant handling to scale without adding coordinators for repetitive data entry
Before the change, each inbound application required someone to open the message, read the resume, and key a contact into the CRM. The exact application volume and labor impact remain, but the workflow constraint was clear: every new business unit added more manual handling.
- Ownership had to be correct before parsing. The recipient address determined the business unit and contact owner. A valid resume attached to the wrong tenant would still be a failed result.
- Applicants controlled the input format. The system had to accept PDF, DOCX, DOC, RTF, HTML, ODT, and TXT instead of asking recruiters to normalize files first.
- The CRM remained the authority. Model output couldn't bypass permissions, validations, initial statuses, or file-attachment behavior already enforced by the application.
Indeed and LinkedIn were sources of application emails, not API integrations. The platform recognized their messages by sender and subject, found the resume link in the email body, and downloaded the document over HTTP. That distinction matters: it avoids implying a partnership or API capability that wasn’t part of the build. This focused case extends Teamvoy’s broader multi-tenant CRM portfolio case. The parent page explains the platform. This one explains the parser that had to work inside it.
03. Solution
The First Production Route Took Roughly Three Weeks
The system follows one ordered route from email receipt to contact creation. Context comes first, extraction second, model output third, and the CRM write last. That order prevents the model from deciding tenancy, ownership, permissions, or what counts as a valid contact.
Email Context Before Model Context
Mailgun inbound routing sends parsed message fields to a Rails endpoint, so the application doesn’t parse raw MIME. The recipient address selects the business unit and owner. Sender and subject identify the source. The endpoint then finds the resume download link and, for Indeed messages, unwraps the redirect before download. Rails answers Mailgun quickly and hands the slower work to Sidekiq. Downloading a file, extracting its contents, and calling a model can take tens of seconds. Moving those steps into a worker protects webhook latency and gives each application a traceable execution path.

– The recipient address sets the business unit and owner before resume data reaches the CRM contact operation.
Mixed-Format Resume Extraction Without a Universal Converter
Mixed-format resume extraction uses a small tool for each document type. Machine-readable PDFs go through pdf-reader. DOCX files use rubyzip and nokogiri to read Word XML. Legacy DOC files use antiword, RTF uses ruby-rtf, HTML uses nokogiri, ODT uses odt2txt, and TXT follows a plain file-read path.
| Input | Production path | Why it exists |
| PDF with a text layer | pdf-reader | Keeps common PDF extraction local to the Rails application |
| PDF without a usable text layer | OpenAI input_file | Gives the model the PDF file rather than pretending empty text is valid input |
| Structurally corrupt PDF | Ghostscript pdfwrite, then re-extraction | Repairs malformed structure before the document is rejected |
| DOCX, DOC, and RTF | Dedicated local tools | Covers current and legacy office formats without an external conversion service |
| HTML, ODT, and TXT | Lightweight format-specific paths | Avoids unnecessary conversion and keeps failures easy to locate |
The trade-off is maintenance. Several narrow extraction paths require more tests than one universal converter. They also avoid another service boundary and keep resume content inside the application environment until the OpenAI request.
OpenAI Structured Outputs as a CRM Contract
The Responses API call uses gpt-4o-2024-08-06 and OpenAI Structured Outputs. A strict JSON Schema defines the ResumeData object, sets strict: true, and rejects undeclared fields through additionalProperties: false. The result is shaped for the contact model rather than returned as prose that the application must interpret afterward. Resume parsing with JSON Schema is different from asking for valid JSON. JSON mode can produce syntactically correct output that still carries the wrong keys or nesting. A strict schema narrows the response to the fields the application knows how to persist. The exact Ruby parameter nesting remains [VERIFY] before any production request example is published. OpenAI’s current documentation describes both Structured Outputs and PDF file inputs. This case does not make claims about retention, model training, residency, or compliance because the project-specific settings have not been approved for publication.
The Existing Rails Operation Creates the Contact
The parser does not insert a separate database record and reconcile it later. It calls the same contact-creation operation used by the CRM interface. Existing validations, permissions, business-unit ownership, initial pipeline status, and upload behavior therefore stay in force. The original resume is attached to the contact and stored in Amazon S3. Keeping the feature inside Rails gave the platform one definition of a valid contact. Teamvoy’s AI Integration Services cover the model boundary, while its Ruby on Rails Development Services map to the application and domain integration.
Technology Used in Production
| Layer | What shipped | Decision behind it |
| Application | Ruby 3.4 and Rails 8.1 | Reused existing models, permissions, business units, uploads, and domain operations |
| Email ingestion | Mailgun inbound routing | Used the platform’s existing provider and supplied parsed fields instead of raw MIME |
| Background work | Sidekiq 8 on Redis | Kept downloads, extraction, and model latency outside the webhook response |
| Document processing | pdf-reader, rubyzip, nokogiri, antiword, ruby-rtf, odt2txt | Covered seven formats without an external extraction platform |
| PDF recovery | Ghostscript pdfwrite | Repaired malformed PDFs before one controlled re-extraction attempt |
| AI parsing | OpenAI Responses API and strict JSON Schema | Produced a predictable ResumeData contract for the CRM |
| Data and files | PostgreSQL, Amazon S3, CarrierWave/fog | Stored contact data and retained the original resume in the existing platform |
| Delivery and monitoring | AWS, ECR, Docker, GitLab CI, Sentry | Ran and observed the feature inside the platform’s production boundary |
The team moved from first implementation work to production hardening in roughly three weeks. Refinement continued in later months, but it extended the same route rather than replacing it. A more detailed phase breakdown remains [VERIFY], so the case does not assign unsupported durations to discovery, testing, or rollout.

– The first production route took roughly three weeks; later months refined the same workflow rather than launching a second system.
The scope stayed narrow: no separate OCR service, new CRM persistence layer, or Indeed and LinkedIn APIs. Each omission reduced the places where ownership, retries, and data handling could diverge.
04. Results
The verified outcome is a live production capability: successful applications move from email to a structured CRM contact without review or confirmation
The system supports seven resume formats and reached its initial hardened production state in roughly three weeks. Performance and business-impact figures remain pending measurement and approval.
| Evidence | Before | Production state | Publication status |
| Format coverage | Applicant files arrived in arbitrary formats | Seven supported formats with explicit extraction paths | Verified |
| Contact creation | A person read the resume and keyed the contact | Successful jobs create contacts through existing CRM logic | Verified |
| Initial delivery | No automated production route | Production hardening in roughly three weeks | Verified |
| Application volume | Manual volume [VERIFY] | Weekly or monthly automated volume [VERIFY] | Needs window |
| Email-to-CRM time | Manual baseline [VERIFY] | Automated median [VERIFY] | Needs method and window |
| Parsing reliability | No supplied baseline | Success and manual-follow-up rates [VERIFY] | Needs sample and window |
The shift is operational rather than hypothetical. Routine applications no longer require a confirmation screen. Work moves to exceptions: expired links, unreadable files, failed extraction, or another logged error. The exception rate must be measured before the page claims saved hours, lower cost, or a percentage improvement.
Duplicate Prevention Over Automatic Recovery
The most important trade-off sits in the retry policy. Sidekiq does not replay the whole job because a partial success followed by a retry could create a duplicate contact. Retries stay narrow around the OpenAI request: two attempts in application code plus the client’s configured retry behavior.

– Whole-job retries stay disabled because replay after a partial success could create a duplicate CRM contact
That policy makes some failures more expensive to recover. A person has to inspect them. It also keeps duplicate risk visible instead of hiding it behind generic queue behavior. A safer automatic recovery path would require idempotency around the contact-creation operation, not a larger retry count.
Four Facts That Define the Build
| Fact | Verified state |
| Production status | Live with no feature flag |
| Document coverage | PDF, DOCX, DOC, RTF, HTML, ODT, and TXT |
| Successful path | Email to CRM contact without review or confirmation |
| Failure path | Logged for manual follow-up; whole-job replay disabled |
05. Conclusion
The strongest part of this AI resume parser was not the prompt. It was the boundary around the model
Email routing established context, local tools handled document variance, OpenAI Structured Outputs enforced the data shape, and the existing Rails operation preserved CRM rules. That combination turned a parsing call into a production workflow.
Talk to Teamvoy About AI Resume Parsing
Tell us how applications enter your system, which document formats arrive in practice, and where the structured data must land.