Direct answer
What dealership leaders need to know
A strong dealership AI pilot uses days 1–30 to baseline one workflow and test in shadow mode, days 31–60 to run a controlled production trial with human review, and days 61–90 to validate economics, controls, adoption, and a stop-or-scale decision. Define the owner, data, permissions, metrics, risks, and exit conditions before day one.- Pilot one operational constraint, not an enterprise-wide promise.
- Baseline the dealership outcome before introducing the AI workflow.
- At day 90, scale only when value, control, and adoption are all supported by evidence.
A 30-60-90 day dealership AI pilot should answer one decision: does this specific workflow create enough safe, measurable operating value to scale? The first 30 days establish the baseline and test the system without broad authority. Days 31–60 run a controlled production pilot with human review. Days 61–90 test repeatability, full cost, controls, adoption, and the case to stop, correct, hold, or scale.
The clock is not the objective. A pilot may need more or less time depending on volume, data quality, risk, seasonality, integration, and the dealership’s ability to observe outcomes. The point of the structure is to prevent a product trial from being mistaken for operating evidence.
This plan is operational guidance, not legal, regulatory, privacy, cybersecurity, employment, or financial advice. Use qualified advisers for the dealership, jurisdiction, data, and workflow involved.
Before day one: write the pilot charter
Do not start with logins and training. Start with a one-page charter that leadership, the workflow owner, and the vendor can interpret the same way.
Complete these fields:
| Charter field | Decision required |
|---|---|
| Constraint | The current workflow problem and who experiences it |
| Scope | Store, department, channel, team, records, and exclusions |
| Outcome | One primary dealership result expected to move |
| Baseline | Current value, calculation, source, and measurement period |
| Leading indicators | Early signals that can be reviewed weekly |
| Risk signals | Errors, complaints, exceptions, access, or compliance indicators |
| Owner | One person accountable for operating performance |
| Technical owner | One person accountable for access, integration, and recovery |
| Executive sponsor | Person who resolves cross-functional constraints and approves scale |
| Permissions | What the system can read, recommend, write, send, or change |
| Stop conditions | Events or thresholds that pause the pilot |
| Decision date | When evidence will be reviewed and what choices are available |
“Improve efficiency with AI” is not a charter. “Reduce the median time to review qualified missed service calls at Store A, without increasing complaints or opt-outs” is narrow enough to test.
Use the AI operator decision framework for car dealerships to select a constraint. If source systems or definitions are unreliable, work through the dealership AI data-readiness guide first.
Choose a pilot with enough signal
A good first workflow is frequent enough to generate evidence, bounded enough to control, reversible if something fails, and important enough that improvement matters. It also has a person willing and able to own it.
Possible candidates include:
- classifying missed fixed-ops calls for human follow-up;
- preparing a daily exception list for unworked leads;
- drafting declined-service follow-up for advisor or BDC review;
- identifying inventory records with missing or inconsistent content;
- summarizing customer history before a scheduled appointment; or
- preparing a management variance report with source links.
Avoid a first pilot that requires broad write access, crosses many departments, lacks a stable source of truth, or creates high-consequence customer representations. The fastest route to evidence is usually a smaller permission surface.
Days 1–30: map, baseline, and test in shadow mode
The first month is for understanding the work and proving that the proposed system can interpret it. It is not a race to switch on automation.
Week 1: observe the existing workflow
Map the actual path, including rework and exceptions:
- What event creates the work?
- Which systems and fields are consulted?
- Which decisions do employees make?
- Where does work wait, duplicate, or disappear?
- Which cases require manager judgment?
- What downstream record or customer experience changes?
Interview the people doing the work. Productive friction often lives outside the written process—in duplicate records, missing notes, stale integrations, approval queues, and local workarounds.
Week 2: establish the baseline
Choose a representative pre-pilot period and record:
- eligible workflow volume;
- completion and cycle time;
- conversion or recovery outcome where applicable;
- employee handling and review time;
- exception, correction, and rework rate;
- customer complaint and opt-out signals;
- technology and labor cost; and
- data completeness and freshness.
Write the formula for every metric. If “response time” starts at lead creation in one report and first employee view in another, the pilot can appear successful without changing the customer experience.
Week 3: map data and permissions
List each source, field, owner, freshness expectation, access method, retention period, and downstream action. Separate permissions into read, recommend, approve, send, and write.
Begin with the lowest useful tier. An agent that prepares a recommendation for a person to review is easier to inspect than one that edits records or contacts customers automatically. The AI agents for car dealerships operating model provides permission tiers and an escalation matrix.
Before DMS, CRM, call, inventory, or customer-data access, use the dealer AI vendor scorecard. The FTC’s automobile dealer Safeguards Rule FAQs are also a relevant primary source for covered dealers evaluating access controls, service providers, monitoring, encryption, multifactor authentication, training, and incident response.
Week 4: run shadow and adversarial tests
Use historical, sandbox, or parallel data without permitting production action. Compare the proposed output with the result an authorized employee would produce.
Test:
- normal and high-volume periods;
- missing, stale, duplicate, and conflicting records;
- unusual language and unsupported requests;
- opt-out and communication preferences;
- unavailable systems and expired access;
- untrusted instructions in documents or notes;
- employee rejection, correction, and escalation; and
- disablement, rollback, and manual recovery.
Review false positives and false negatives separately. An average quality score can hide the failure type that matters most.
Day-30 gate
Move to controlled production only if the team can show:
- an agreed baseline;
- a complete workflow and data map;
- named owners and trained reviewers;
- acceptable test performance for the use case;
- explicit permissions and prohibited actions;
- working escalation, logging, stop, and fallback paths; and
- no unresolved critical vendor or security issue.
If those conditions are missing, extend or stop. Calendar progress is not readiness.
Days 31–60: run controlled production
The second month tests the workflow with real operating conditions and limited authority. Keep the sample bounded by store, team, channel, hours, or eligible case type.
Start with human review
Reviewers should see the source facts, proposed action, uncertainty or exception, and ability to edit or reject. Capture the decision and reason. That feedback becomes operating evidence, not merely model feedback.
Train employees on:
- what the system does and does not do;
- which source remains authoritative;
- how to review and correct an output;
- when and how to escalate;
- how to report an incident or unexpected behavior;
- what is logged; and
- how pilot performance affects their work.
NADA’s dealer education has emphasized practical AI adoption and employee understanding, including its discussion of using AI in the dealership. The dealership AI change-management framework turns that principle into role, training, and adoption decisions.
Hold a weekly operating review
Use the same agenda every week:
- primary outcome versus baseline;
- eligible volume, coverage, and exclusions;
- completion, edit, rejection, override, and escalation rates;
- customer, employee, data, and system risk signals;
- aged exceptions and unresolved incidents;
- work saved, added, or shifted to another role;
- vendor or integration changes; and
- continue, correct, restrict, or pause decisions.
Inspect a small sample of successes, failures, and near misses. Dashboards show frequency; case review reveals mechanism.
Use a balanced pilot scorecard
| Dimension | Example measure | Guardrail |
|---|---|---|
| Outcome | Qualified appointment show rate | Same definition and source as baseline |
| Speed | Median eligible-case cycle time | Report distribution, not only average |
| Quality | Correct completion after review | Segment material and minor corrections |
| Adoption | Eligible cases reviewed through workflow | Do not punish justified overrides |
| Customer | Complaint, correction, and opt-out rate | Compare with baseline and channel volume |
| Control | Unauthorized action attempts | Critical events can be immediate stop conditions |
| Economics | Net monthly benefit at observed volume | Include review, setup, integration, and management time |
Activity metrics such as summaries generated, conversations handled, or messages sent can explain operation. They do not prove value.
Day-60 gate
Continue only when the workflow is stable enough to evaluate, employees use it as designed, exceptions are owned, critical controls work, and preliminary outcome movement is not being purchased with unacceptable risk or hidden labor.
Do not expand scope and permissions at the same time. Change one meaningful variable, then observe its effect.
Days 61–90: validate repeatability and economics
The third month tests whether the result survives normal variation and whether the dealership can operate the workflow without constant rescue.
Reconcile the economics
Include:
- software and usage fees;
- implementation and integration cost;
- data cleanup;
- employee training;
- human review and exception handling;
- management and reporting time;
- security, compliance, and procurement effort;
- vendor support and workflow maintenance; and
- switching or exit cost.
Document the value formula, volume, attribution, time horizon, and uncertainty. Use the dealership AI ROI calculator and measurement framework to avoid treating gross activity as net benefit.
Test operational independence
Ask whether dealership employees can:
- explain the workflow and boundaries;
- find the source behind an output;
- resolve common exceptions;
- revoke access and stop actions;
- operate the manual fallback;
- retrieve logs and performance data; and
- challenge a vendor-reported result.
If only the vendor understands the system, the dealership has a dependency—not an operating capability.
Segment the result
An aggregate can conceal where the pilot works. Review performance by store, team, channel, source, case type, time period, and relevant customer journey stage. Note sample size and avoid declaring a durable pattern from a small slice.
The NIST AI Risk Management Framework Core treats AI risk management as a continuous govern, map, measure, and manage process. A day-90 approval should therefore define the next monitoring cycle, not close the file.
Make one of four decisions at day 90
Scale
The outcome improved, the economics remain credible, controls worked, employees adopted the workflow, and the team can operate it. Expand one dimension at a time with a new baseline and owner.
Hold
Evidence is promising but the sample, season, integration, or outcome window is insufficient. Continue within the same boundaries and set a new decision date.
Correct
The use case remains valid, but data, workflow, product behavior, training, or ownership prevented a fair result. Write a correction plan and retest rather than quietly redefining success.
Stop
Value is weak, full cost is too high, risks are unacceptable, the vendor cannot support required controls, or ownership is absent. Revoke access, preserve required evidence, export needed records, confirm deletion obligations, restore the fallback, and document the lesson.
A stopped pilot can be an excellent decision. It prevents an unproven tool from becoming permanent infrastructure.
Stop conditions to define in advance
Immediate or threshold-based conditions may include:
- unauthorized message, record change, permission, or system action;
- material customer-facing factual error;
- mishandled opt-out or communication preference;
- unexplained customer or dealership data access;
- security event without effective containment and notification;
- missing audit trail for consequential actions;
- sustained rise in complaints, corrections, or employee workarounds;
- integration failure that makes outputs stale or misleading;
- vendor change that materially alters data use or behavior; or
- loss of the named operating or technical owner.
Write who can pause the pilot and how. A control that requires a committee meeting during an active incident is not a usable stop control.
The day-90 evidence pack
Keep the final review short enough to use and specific enough to audit:
- charter and original baseline;
- workflow, data, and permission maps;
- vendor and control decisions;
- training and adoption record;
- scorecard with definitions and sources;
- exception, incident, correction, and override analysis;
- full-cost economic model;
- scale, hold, correct, or stop recommendation;
- accountable owner and next review date; and
- access, retention, monitoring, and exit plan.
Take the Dealership AI Operational Depth assessment before and after the pilot to identify whether workflow, data, human ownership, controls, or evidence is the limiting pillar. The assessment is directional while dealer data collection continues; it does not currently publish a peer percentile or representative benchmark.
The practical standard
At day 90, the dealership should be able to explain what changed in the operation, what it cost, which risks appeared, how people retained control, and why the evidence supports the next decision. If the conclusion relies mainly on a vendor dashboard, a successful demo, or employee enthusiasm, the pilot has not yet answered the management question.
The goal is not to prove that AI works. It is to determine whether this AI workflow works here, under dealership-owned conditions, with enough value and control to deserve its next level of permission.