FIXED SCOPE
AI & System Readiness Audit

Architecture review, risk surface, prioritised action plan. No obligation.

PAID - 2 WEEKS
Sharp Sprint

Fixed scope, senior engineers, working software. Skip the long discovery.

Contact us

AI Resume Parser: Seven Formats, One CRM Workflow

An AI resume parser moved applicant data from inbound email into an existing Rails CRM without adding a separate review screen. It handled seven document formats, malformed PDFs, scanned files, and source-specific links inside one production route. How do you automate that workflow without bypassing the rules that already make a CRM contact valid?

See how a production AI resume parser turns application emails and seven resume formats into structured contacts inside an existing Rails CRM.
See how a production AI resume parser turns application emails and seven resume formats into structured contacts inside an existing Rails CRM.
See how a production AI resume parser turns application emails and seven resume formats into structured contacts inside an existing Rails CRM.
See how a production AI resume parser turns application emails and seven resume formats into structured contacts inside an existing Rails CRM.
ui screenshot of an applicant crm dashboard featuring a resume parser, with a left navigation menu (inbound, parser, contacts) and a main workspace titled'Resume parser' on a monitor frame.
Project detailValue
ClientAn US insurance marketing platform
IndustryInsurance
StatusLive in production; no feature flag
Initial deliveryRoughly three weeks to production hardening
RefinementContinued in later months
Core stackRuby 3.4, Rails 8.1, Sidekiq 8, OpenAI Responses API, PostgreSQL, AWS
three rounded cards with large numbers: ~3, 7, and 2. card 1 notes ~3 weeks to the initial production path; card 2 lists 7 supported resume formats; card 3 shows 2 input branches for model-ready content.

01. Summary

The AI resume parser converts an application email into a structured contact inside the client’s CRM

Mailgun receives the message, Rails determines the business unit and owner, Sidekiq handles the slow work, format-specific tools extract the resume, and OpenAI Structured Outputs returns data that matches the CRM schema. A parser demo usually starts with a clean PDF and ends with JSON. This project started earlier and ended later. Applications arrived as email notifications from Indeed and LinkedIn. The resume could sit behind an expiring link, use a legacy document format, contain no usable text layer, or fail ordinary PDF extraction. The output then had to pass the same validations and permissions as a contact created by a person. The model call wasn’t the product. The production work sat around it: source detection, ownership, document recovery, schema enforcement, CRM domain logic, and failure handling. Successful jobs run unattended. Failed jobs are logged for manual follow-up because replaying the whole job could create a duplicate contact.

02. Problem

The recruiting business units needed applicant handling to scale without adding coordinators for repetitive data entry

Before the change, each inbound application required someone to open the message, read the resume, and key a contact into the CRM. The exact application volume and labor impact remain, but the workflow constraint was clear: every new business unit added more manual handling.

Three requirements made this more than a document-parsing feature:
  • Ownership had to be correct before parsing. The recipient address determined the business unit and contact owner. A valid resume attached to the wrong tenant would still be a failed result.
  • Applicants controlled the input format. The system had to accept PDF, DOCX, DOC, RTF, HTML, ODT, and TXT instead of asking recruiters to normalize files first.
  • The CRM remained the authority. Model output couldn't bypass permissions, validations, initial statuses, or file-attachment behavior already enforced by the application.

Indeed and LinkedIn were sources of application emails, not API integrations. The platform recognized their messages by sender and subject, found the resume link in the email body, and downloaded the document over HTTP. That distinction matters: it avoids implying a partnership or API capability that wasn’t part of the build. This focused case extends Teamvoy’s broader multi-tenant CRM portfolio case. The parent page explains the platform. This one explains the parser that had to work inside it.

03. Solution

The First Production Route Took Roughly Three Weeks

The system follows one ordered route from email receipt to contact creation. Context comes first, extraction second, model output third, and the CRM write last. That order prevents the model from deciding tenancy, ownership, permissions, or what counts as a valid contact.

Email Context Before Model Context

Mailgun inbound routing sends parsed message fields to a Rails endpoint, so the application doesn’t parse raw MIME. The recipient address selects the business unit and owner. Sender and subject identify the source. The endpoint then finds the resume download link and, for Indeed messages, unwraps the redirect before download. Rails answers Mailgun quickly and hands the slower work to Sidekiq. Downloading a file, extracting its contents, and calling a model can take tens of seconds. Moving those steps into a worker protects webhook latency and gives each application a traceable execution path.

hero section explaining that email context reaches the crm before model output, with a workflow diagram of email-to-crm data flow.

Mixed-Format Resume Extraction Without a Universal Converter

Mixed-format resume extraction uses a small tool for each document type. Machine-readable PDFs go through pdf-reader. DOCX files use rubyzip and nokogiri to read Word XML. Legacy DOC files use antiword, RTF uses ruby-rtf, HTML uses nokogiri, ODT uses odt2txt, and TXT follows a plain file-read path.

InputProduction pathWhy it exists
PDF with a text layerpdf-readerKeeps common PDF extraction local to the Rails application
PDF without a usable text layerOpenAI input_fileGives the model the PDF file rather than pretending empty text is valid input
Structurally corrupt PDFGhostscript pdfwrite, then re-extractionRepairs malformed structure before the document is rejected
DOCX, DOC, and RTFDedicated local toolsCovers current and legacy office formats without an external conversion service
HTML, ODT, and TXTLightweight format-specific pathsAvoids unnecessary conversion and keeps failures easy to locate

The trade-off is maintenance. Several narrow extraction paths require more tests than one universal converter. They also avoid another service boundary and keep resume content inside the application environment until the OpenAI request.

OpenAI Structured Outputs as a CRM Contract

The Responses API call uses gpt-4o-2024-08-06 and OpenAI Structured Outputs. A strict JSON Schema defines the ResumeData object, sets strict: true, and rejects undeclared fields through additionalProperties: false. The result is shaped for the contact model rather than returned as prose that the application must interpret afterward. Resume parsing with JSON Schema is different from asking for valid JSON. JSON mode can produce syntactically correct output that still carries the wrong keys or nesting. A strict schema narrows the response to the fields the application knows how to persist. The exact Ruby parameter nesting remains [VERIFY] before any production request example is published. OpenAI’s current documentation describes both Structured Outputs and PDF file inputs. This case does not make claims about retention, model training, residency, or compliance because the project-specific settings have not been approved for publication.

The Existing Rails Operation Creates the Contact

The parser does not insert a separate database record and reconcile it later. It calls the same contact-creation operation used by the CRM interface. Existing validations, permissions, business-unit ownership, initial pipeline status, and upload behavior therefore stay in force. The original resume is attached to the contact and stored in Amazon S3. Keeping the feature inside Rails gave the platform one definition of a valid contact. Teamvoy’s AI Integration Services cover the model boundary, while its Ruby on Rails Development Services map to the application and domain integration.

Technology Used in Production

LayerWhat shippedDecision behind it
ApplicationRuby 3.4 and Rails 8.1Reused existing models, permissions, business units, uploads, and domain operations
Email ingestionMailgun inbound routingUsed the platform’s existing provider and supplied parsed fields instead of raw MIME
Background workSidekiq 8 on RedisKept downloads, extraction, and model latency outside the webhook response
Document processingpdf-reader, rubyzip, nokogiri, antiword, ruby-rtf, odt2txtCovered seven formats without an external extraction platform
PDF recoveryGhostscript pdfwriteRepaired malformed PDFs before one controlled re-extraction attempt
AI parsingOpenAI Responses API and strict JSON SchemaProduced a predictable ResumeData contract for the CRM
Data and filesPostgreSQL, Amazon S3, CarrierWave/fogStored contact data and retained the original resume in the existing platform
Delivery and monitoringAWS, ECR, Docker, GitLab CI, SentryRan and observed the feature inside the platform’s production boundary

The team moved from first implementation work to production hardening in roughly three weeks. Refinement continued in later months, but it extended the same route rather than replacing it. A more detailed phase breakdown remains [VERIFY], so the case does not assign unsupported durations to discovery, testing, or rollout.

two black rounded panels showing project phases: left panel reads'~3 weeks / Initial production path / implementation through production hardening'; right panel reads 'Later months / Production refinement / continued after the first live path'.

The scope stayed narrow: no separate OCR service, new CRM persistence layer, or Indeed and LinkedIn APIs. Each omission reduced the places where ownership, retries, and data handling could diverge.

04. Results

The verified outcome is a live production capability: successful applications move from email to a structured CRM contact without review or confirmation

The system supports seven resume formats and reached its initial hardened production state in roughly three weeks. Performance and business-impact figures remain pending measurement and approval.

EvidenceBeforeProduction statePublication status
Format coverageApplicant files arrived in arbitrary formatsSeven supported formats with explicit extraction pathsVerified
Contact creationA person read the resume and keyed the contactSuccessful jobs create contacts through existing CRM logicVerified
Initial deliveryNo automated production routeProduction hardening in roughly three weeksVerified
Application volumeManual volume [VERIFY]Weekly or monthly automated volume [VERIFY]Needs window
Email-to-CRM timeManual baseline [VERIFY]Automated median [VERIFY]Needs method and window
Parsing reliabilityNo supplied baselineSuccess and manual-follow-up rates [VERIFY]Needs sample and window

The shift is operational rather than hypothetical. Routine applications no longer require a confirmation screen. Work moves to exceptions: expired links, unreadable files, failed extraction, or another logged error. The exception rate must be measured before the page claims saved hours, lower cost, or a percentage improvement.

Duplicate Prevention Over Automatic Recovery

The most important trade-off sits in the retry policy. Sidekiq does not replay the whole job because a partial success followed by a retry could create a duplicate contact. Retries stay narrow around the OpenAI request: two attempts in application code plus the client’s configured retry behavior.

two-part image: left shows text about duplicate prevention and automatic recovery; right shows a dashed bordered panel titled'Controlled Job Boundary' with a flow: Download → Extract → OpenAI request → CRM operation.

That policy makes some failures more expensive to recover. A person has to inspect them. It also keeps duplicate risk visible instead of hiding it behind generic queue behavior. A safer automatic recovery path would require idempotency around the contact-creation operation, not a larger retry count.

Four Facts That Define the Build

FactVerified state
Production statusLive with no feature flag
Document coveragePDF, DOCX, DOC, RTF, HTML, ODT, and TXT
Successful pathEmail to CRM contact without review or confirmation
Failure pathLogged for manual follow-up; whole-job replay disabled

05. Conclusion

The strongest part of this AI resume parser was not the prompt. It was the boundary around the model

Email routing established context, local tools handled document variance, OpenAI Structured Outputs enforced the data shape, and the existing Rails operation preserved CRM rules. That combination turned a parsing call into a production workflow.

The fastest way in: book a 15-minute call with a Chief Technology Officer this week.
PREFER email?
Teamvoy can map the ingestion, schema, domain-operation, and failure boundaries
before your team commits to a model or queue design.
cropped-avatar
Bohdan Varshchuk
Chief Technology Officer

Start the conversation with Teamvoy