Data Mapping for SaaS Implementations: Playbook + Template
Data mapping is the written agreement of which source field lands in which target field, with what transformation, what default when blank, and who owns the decision. Do it during discovery, one 60 to 90 minute workshop per object, anchored to 20 to 50 real records rather than a schema printout. Most rework comes from six decisions made late: picklist values, dates and time zones, identifiers, field lengths, required fields, and blanks and duplicates. Keep one versioned sheet per object with a changelog, get a named owner on each side to sign each version, and prove it with a small test load before the first full dry run.
Data mapping is the written agreement of which field in the customer's source system lands in which field of your product, with what transformation, what default when the value is blank, and who owns the decision. In a SaaS implementation it is the one document that turns "migrate our data" into work an engineer can build and a customer can sign. Do it during discovery, before configuration, one workshop per object, anchored to real records instead of a schema printout.
It matters because migrations rarely fail on the transfer. They fail on meaning: a status value that does not exist in the target, a date that shifts by a day, an account number that lost its leading zero. An analysis of more than 500 enterprise migration reviews found poor discovery and planning behind 68% of negative outcomes and data quality or integrity failures behind 54%. This guide covers what the mapping document needs, how to run the workshop, the decisions that cause the most rework, who owns what, and how to test before you load anything.
What is data mapping in a SaaS implementation?
Data mapping is the process of establishing the relationship between each source field and its target, including the transformation rule that gets one to the other. Future Processing defines source-to-target mapping as identifying the source systems, understanding their structure and content, and determining how each element maps or transforms into the target. In a customer onboarding, the source is whatever the customer runs today: a legacy CRM, an ERP export, a dozen spreadsheets, or a competitor's product. The target is your data model.
It happens at three levels, and the order matters:
- Object mapping. Their "clients" are your "accounts"; their "jobs" are your "projects". Decide the object correspondence first, including one-to-many cases such as a source contact that belongs to two companies, before touching a single field.
- Field mapping. For each target field, which source field feeds it, a fixed default, or nothing.
- Value mapping. For enumerated fields, which source value becomes which target value. "In progress", "WIP", and "Open" may all need to become "Active".
Mapping is one workstream inside the wider customer data migration, and it depends on process discovery having happened first. Discovery tells you how the customer works; mapping records how their data will be represented in your product. Integration requirements ask the same field-level questions for systems that stay connected after go-live, where the mapping has to hold every day rather than once.
Why does data mapping decide whether the migration lands?
Because the calendar time goes to mapping and validation rather than to moving bytes, and because the errors that hurt are semantic. A research study Experian commissioned with Data Migration Pro found that only 36% of migration projects kept to their original budget and 54% suffered some form of delay. Kanerika's 2026 analysis of more than 500 enterprise reviews on G2, Gartner Peer Insights, and TrustRadius attributes 68% of negative reviews to poor discovery and planning and 54% to data quality and integrity failures, with one retail enterprise reporting that it migrated 18 months of records only to find schema mismatches it had never mapped.
The cost lands after go-live. Gartner puts the cost of poor data quality at an average of $12.9 million a year per organization. In an implementation the same problem shows up smaller and sooner: a picklist that imported wrong means every report the customer builds in week one is wrong, and every fix arrives as a support ticket against your team.
The pattern behind the numbers is that teams map from the schema instead of the data. A schema tells you a field exists and what type it is. It does not tell you that "Region" has 14 values in production, 9 of them typed by hand, or that "Close date" is blank on 40% of records because sales stopped filling it in two years ago. Only the records tell you that, which is why the workshop below starts with a sample.
When should you map, and how long does it take?
Start mapping in discovery, before configuration, and finish version 1 before you build anything that depends on the data model. Two reasons. First, mapping surfaces configuration decisions: if the customer has 14 regions and your product ships with 6, you either add values or collapse them, and that choice touches every screen and report. Second, mapping produces the first real list of customer deliverables: the exports, the value lists, and the decisions only they can make.
Effort is driven by the number of objects and enumerated fields, not by row count. These are working rules of thumb from running these sessions, not benchmarks:
| Migration shape | Objects | Mapping effort | Elapsed time |
|---|---|---|---|
| Single source, standard objects (accounts, contacts, a few custom fields) | 2 to 4 | One 90-minute workshop plus 2 to 4 hours of follow-up | About 1 week |
| Single source with custom objects and history | 5 to 10 | One workshop per object, 60 to 90 minutes each | 2 to 3 weeks |
| Multiple sources merging into one model | 10 or more | Object-level reconciliation first, then per-object workshops, plus a dedup rule per object | 4 to 6 weeks, often the critical path |
Elapsed time is dominated by waiting: for exports, for a value list, for the one person who knows what "Type 3" meant in 2019. Put the customer-side tasks on the deliverable register with dates the day the workshop ends.
What should a data mapping document contain?
One workbook per migration, one sheet per object, plus a changelog and a transformations tab. Future Processing recommends exactly that layout, a changelog, a notes tab, a mapping tab per object, and a separate sheet for complex transformations, and it survives contact with real projects because each object sheet can be versioned and signed on its own.
Columns for the per-object sheet:
| Column | What goes in it | Why it matters |
|---|---|---|
| Source object and field | The exact API or column name, not the screen label | Labels get renamed; API names are what the export contains |
| Source type and length | text(255), picklist, date, datetime, number(10,2) | Type and length mismatches get found here instead of on load day |
| Sample values | 3 to 5 real values from production | Exposes hand-typed variants, blanks, and legacy codes |
| Fill rate | Percent of records with a value | A 12% fill rate changes whether the field is worth migrating at all |
| Target object and field | Your API name | Same reason as the source column |
| Transformation rule | Direct, value map (see tab), concatenate, split, truncate, derive | The rule is what the engineer builds and what the customer signs |
| Default when blank | A fixed value, leave empty, or reject the record | Blank handling differs by tool, so decide it explicitly |
| Required in target | Yes or no, plus any validation rule | Required fields with low fill rates are the top cause of failed rows |
| Decision owner | A named person on the customer side | Mapping questions stall when they are addressed to "the customer" |
| Status | Proposed, agreed, built, tested, signed | Lets you report percent agreed per object in the status update |
| Source of decision | Call date, Slack link, or email | Someone will ask in month three why a rule exists |
Two habits keep the workbook usable. Freeze the left side: once the source export is agreed, never reorder or rename it; put every change on the target side and log it in the changelog with a date and who asked for it. And version by object, so the accounts sheet can be at v3 and signed while contacts is still at v1.
How do you run a data mapping workshop?
Prepare three things, run 60 to 90 minutes per object, and leave with a signed v1 plus an open-decisions list. Do not run it from the schema alone.
Before the session
- A production export of 20 to 50 real records per object, with personal data masked where required. Test data defeats the purpose; the point is to see the actual variants.
- A value inventory for every picklist, status, and category field: each distinct value and its record count. It is a five-minute query for the customer's admin, and it settles half the arguments before they start.
- Your target field list with type, length, required flag, and allowed values, pre-filled into the right-hand side of the sheet with your proposed mapping. The workshop confirms and corrects; it should never start from blank.
Who attends
On the customer side: the admin who can export and explain the schema, and one person who uses the data daily and knows what the values mean in practice. On your side: the implementer who owns the plan and whoever will build the load. Four people is right. Eight is a status meeting.
An agenda that works
- Object confirmation (5 minutes). Confirm the object correspondence and the record count to migrate per object. Agree what stays behind.
- Walk the target, not the source (40 minutes). Go field by field down your model. For each field: which source field, what rule, what default. Walking the source instead produces a long list of fields nobody needs.
- Value mapping for enumerations (20 minutes). Put the value inventory on screen and map each source value to a target value. Ask about every low-count value; those are the hand-typed ones.
- Decisions and parking lot (10 minutes). Read back every decision made, and give every open one an owner and a date. Get an explicit "agreed" on the sheet.
Anchor every question to a record on screen. "What is Region used for?" gets a policy answer. "Here are the 14 Region values, which of these do you want in the new system?" gets a decision. Once a decision is made, mark it agreed and move on; reopening decisions is what turns two workshops into six.
Which mapping decisions cause the most rework?
Six categories account for most failed loads and post-go-live corrections. Each is cheap to decide in the workshop and expensive to discover afterward.
1. Picklist and enumeration values
Target systems are strict about allowed values. HubSpot's import requires enumeration values to match either an option's label or its internal value, and for default properties the label has to be in English. Salesforce's Data Import Wizard handles a value that does not match an unrestricted picklist by accepting it as a new value, and on a restricted picklist by substituting the default or failing the row. Neither outcome is what you want. Map every value explicitly, including the misspellings, and decide whether unmapped values reject the row or land in a holding value such as "Needs review".
2. Dates, times, and time zones
A date without a time is ambiguous. Salesforce documents that Data Loader converts the date in the file to GMT, so on a machine outside GMT, or during daylight saving time, the imported date can be off by a day. HubSpot sets the time to midnight when a timestamp is omitted and accepts an explicit offset such as +01:00. Record the source format, the source time zone, and the target rule for every date field. A contract end date that shifts by a day will be noticed by the customer's finance team.
3. Identifiers and numeric codes
Any ID that passes through a spreadsheet is at risk. Excel removes leading zeros and keeps a maximum of 15 significant digits, so digits past the 15th are rounded to zero, which silently corrupts a 16-digit account number. The same auto-conversion turned gene names into dates in roughly one-fifth of genomics papers with supplementary Excel gene lists, according to a 2016 scan published in Genome Biology. Keep identifiers as text end to end, and decide which ID is the match key for updates: HubSpot requires a unique identifier such as record ID, email, or company domain in the file to update existing records instead of creating duplicates.
4. Field length and type
Truncation is the quiet one. If the source field holds 16 characters and the target holds 8, you either change the target or agree how to shorten the value; Future Processing calls out exactly that case. Free text heading into a structured field, such as a "Notes" column that someone wants to become a picklist, needs a rule or a decision to leave it in a notes field.
5. Required fields and validation rules
Salesforce runs validation rules on records before import and does not import rows that fail them. A required target field with a 60% fill rate in the source means 40% of rows fail unless you set a default or relax the rule for the load. Decide per field in the workshop instead of finding out in the error log.
6. Blank handling and duplicates
Blank cells behave differently by tool. HubSpot ignores blank cells on import and will not clear an existing value; other tools overwrite. Decide, per field, whether blank means "no change", "clear it", or "reject". For duplicates, agree the dedup rule per object (email, domain, external ID, or a composite) before the first test load, because merging after import is manual work the customer will expect you to do.
Who owns data mapping: your team or the customer's?
Split it by what each side knows. The customer owns the meaning of their data: what each value represents, which records are current, what stays behind, and every business decision (collapse 14 regions to 6, or not). Your team owns the target constraints, the transformation rules, the load mechanics, and the document itself. Name one decision owner per side and put them in the sheet. "The customer" is not an owner.
| Activity | Customer | Your team |
|---|---|---|
| Export the source schema and 20 to 50 sample records per object | Owns | Specifies the format |
| Value inventory for enumerated fields | Provides | Requests and reviews |
| Proposed field mapping (v0) | Reviews | Owns |
| Business decisions on values, defaults, and what to leave behind | Owns | Advises on consequences |
| Transformation rules and load scripts | Signs off | Owns |
| Data cleanup in the source | Owns, unless contracted | Reports what needs cleaning |
| Test load reconciliation and sign-off | Verifies against their own reports | Runs the load and reports counts |
Sign-off is per object and per version, in writing, on the sheet. If your contract charges for data cleanup, the fill-rate and value-inventory columns are also your scope evidence when the customer's "clean" export turns out to have nine ways to spell one state.
How do you test a mapping before the real load?
Load a small file first, reconcile counts per object, then spot-check fields against the customer's own reports. Salesforce's guidance is to import a small test file before the full load, and the same rule applies to every target. Unmapped fields are not imported at all, so the test load also catches anything you forgot.
- Test load of 20 to 50 records per object into a sandbox, using the exact mapping and rules in the sheet. Every row that fails goes back to the sheet as a rule change, never as a manual fix.
- Count reconciliation. Source rows, rows loaded, rows rejected, and why, per object. Rejected rows must equal the rows you chose to leave behind plus rows with a documented failure.
- Field-level spot checks. Pick 10 records the customer knows well and compare every mapped field side by side. Dates, picklists, and IDs first.
- Report parity. Have the customer rebuild one report they rely on (open opportunities by stage, active contracts by region) in the target and compare totals with the source. This is the test that finds value-mapping errors.
- Repeat at full volume in a dry run before cutover, and run at least two of them. The UAT plan should include the customer's report parity check as a scripted scenario.
Mapping v1 is not done until the test load reconciles. Treat the status column as the truth: nothing moves to "signed" without a passing test.
How do you stop mapping decisions from getting lost?
Most mapping decisions are made outside the workshop. They happen on Thursday's call when the admin says "actually, Type 3 was a pilot program, drop those", and in the Slack thread where someone agrees to default the owner to the account manager. If the sheet is only updated by whoever remembers, it drifts from what was agreed, and the dry run fails on rules nobody wrote down.
Three practices hold it together. First, every decision gets a row in the changelog the same day, with its source, which is why the sheet has a "source of decision" column. Second, open mapping questions go on the RAID log or the deliverable register with an owner and a date, so they get chased like any other customer dependency instead of sitting in a workbook tab. Third, capture the conversation itself. This is where Stipulate fits for implementation teams: it reads the customer Slack channel and call transcripts and builds a record of each decision, requirement, and risk linked to the message it came from, so "we agreed to collapse the regions" is findable in month three with its evidence rather than reconstructed from memory. The mapping document remains the contract; the conversation record is how you keep it honest. For the wider habit, see how to track decisions in customer Slack channels.
Next steps
- List the objects in scope and the record count for each. Agree what stays behind before mapping a single field.
- Request a production export of 20 to 50 records per object and a value inventory for every enumerated field. Put both on the deliverable register with dates.
- Build the workbook: one sheet per object with the columns above, a changelog, and a transformations tab. Pre-fill your proposed mapping.
- Run one 60 to 90 minute workshop per object with four people, walking your target model field by field against real records.
- Decide the six rework categories explicitly per field: values, dates and time zones, identifiers, lengths, required fields, blanks and duplicates.
- Get a named owner on each side to sign each object sheet, per version.
- Test-load 20 to 50 records, reconcile counts, spot-check fields, and rebuild one customer report before the first full dry run.
- Log every later decision with its source the day it is made, and track open questions on the RAID log.