| Takeaway | Detail |
|---|---|
| Article 4’s AI-literacy duty applies from February 2, 2025. | The guide treats that date as the compliance starting point for workplace AI literacy under the EU AI Act. |
| Use a 60/40 train-or-verify split for retrieval coaching. | The headline’s 60/40 rule frames coaching around training users and verifying retrieved outputs before commitment. |
| RAG adds external information to LLM responses. | Retrieval-augmented generation retrieves relevant information from external data sources and incorporates it into responses. |
| Verify the live, complete option before committing. | Compare like-for-like totals and terms, and do not rely on an incomplete or outdated option. |
This guide translates the February 2, 2025 Article 4 AI-literacy duty into a practical train-or-verify approach for workplace retrieval coaching.
It explains how RAG uses external information while applying a live-option, like-for-like verification rule before decisions are committed.
How It Works
Retrieval-augmented generation splits a single answer into two operations. Wikipedia defines RAG as a technique that enables large language models to retrieve and incorporate new information from external data sources: the model first refers to a specified set of documents, then responds, supplementing what it learned in training. NVIDIA describes the same pipeline as enhancing accuracy and reliability with facts fetched from external sources. In a workplace coaching tool, this means the assistant is not reciting your policy from memory. It fetches passages from a document set you control, then drafts a suggestion over those passages.
Article 4 attaches a duty to that pipeline. From 2 February 2025, providers and deployers of AI systems must take measures to ensure a sufficient level of AI literacy among their staff and other people who deal with the operation and use of those systems on their behalf. The sufficiency test is relative: it accounts for the person's technical knowledge, experience, education, and training, and for the context in which the system is used. The mechanism is therefore a mapping exercise — role, decisions, comprehension required — framed as measures rather than a prescribed credential.
Five terms carry the mechanism, and each has a matching verification habit:
| Term | What it denotes | Verify by |
|---|---|---|
| Corpus | The specified document set the retriever may draw from | Confirm the policy or agreement version is current |
| Retriever | The component that selects passages at query time | Ask which documents were searched for this answer |
| Retrieved context | The passages handed to the model — the only new facts in play | Open the passage itself, not the summary of it |
| Generation | The drafting step, blending retrieved text with prior training | Separate sourced claims from generic phrasing |
| Deployer | The organization putting the system into use | Identify whose staff the literacy duty covers |
The practical consequence for coaching loops is that the retrieved passage, not the sentence, is the unit of truth. Because generation blends retrieved documents with pre-existing training data, an answer can be partly grounded and partly generic in the same paragraph. The check is two-part: does the output point to a passage, and does that passage actually state what the answer claims? If either half fails, treat the suggestion as unverified and read the source before acting.
The duty reaches the whole chain, not just the end user. The "on their behalf" language covers people operating and using the system, which in a coaching deployment includes whoever curates the corpus as well as the manager receiving a suggestion. Map each role to the artifact it must be able to read — policy text, source citation, or system output — and confirm that mapping covers everyone whose decisions the coaching output can influence.
Key Factors to Consider
The top three decision criteria are AI-literacy coverage, retrieval quality, and total cost under matching terms. For an Article 4 review, verify that the coaching arrangement gives each relevant user enough practical knowledge to use the workplace system responsibly, not merely access to a prompt interface. Ask for the live curriculum, delivery format, audience, completion evidence, and update process. The required level should reflect the users’ technical knowledge, experience, education or training, and the context in which they use the system.
For retrieval quality, check whether the provider identifies the sources available to the coach, explains how source freshness is managed, and gives users a way to inspect supporting material. Wikipedia describes RAG as using specified external documents to supplement a model’s existing training data. That makes source governance a purchasing criterion: verify permitted repositories, access controls, document ownership, citation behavior, and the process for correcting or removing outdated workplace material. Treat a demonstration as evidence only if it uses the same source set and permissions offered in the live option.
For cost, build one total for the same user count, billing period, implementation scope, support level, data connections, storage or usage charges, and AI-literacy delivery. Separate recurring charges from one-time work, and record what happens when the term, seats, support package, or source requirements change. Coworker.ai’s launch coverage says its organizational-memory layer promises to be 9x cheaper; treat that as a supplier claim to validate, not as a transferable saving. Request the baseline, calculation, included services, and comparable terms before relying on it.
Use a written verification sheet before committing. Mark whether the live option includes the complete coaching workflow, the stated training materials, source inspection, administrator controls, user support, and records showing who received the literacy instruction. Then reconcile each item against the order form, service description, privacy terms, and implementation statement. If a feature appears only in a presentation or trial, list it as unconfirmed rather than counting it in the total.
Finally, assess fit by asking what a user must know before acting on a retrieved coaching answer: how to judge source authority, recognize an incomplete response, protect confidential information, and escalate a consequential decision. NTT DATA describes RAG as improving accuracy, contextual understanding, and cost-effectiveness by integrating relevant external information, but those benefits do not replace user judgment. Commit only when the evidence, training scope, and complete commercial total are available for the exact live configuration.
Common Mistakes
The most expensive error in an Article 4 review is verifying a version of the coaching tool you will never actually run. Vendors demonstrate with a curated document set and an administrator login; the contract delivers a production workspace with role-based permissions and whatever repositories happen to be connected on day one. That gap is where AI-literacy claims quietly fail, because a coach that cannot reach the source material still answers — just from the base model's training data.
Consider a compliance team that pilots a retrieval-augmented coaching assistant against a hand-picked folder of HR policies. Answers are accurate, citations are clean, and the team signs. Production then points the same assistant at a shared drive where most employees' permissions exclude those policy documents. For a standard employee account, retrieval returns nothing relevant and the coach improvises. The pre-commit check is concrete: before signature, run your own question set against the live production index while logged in as a standard employee, not an admin, and confirm each answer surfaces a source that user is permitted to open. If the vendor will not provide a live tenant to test, you are verifying a demo, not the option.
A second pitfall is counting seats sold as literacy delivered. Enrollment exports and completion records diverge, and a claim built from assignments will not survive an auditor's question about who actually received practical knowledge. Ask for the completion report you would receive at reporting time, not the enrollment export, and ask which roles the configuration excludes by default — contractors, temporary staff, and workers whose access requests are still pending are the usual omissions.
A third mistake is comparing unlike totals. A quote may cover coaching seats while the retrieval connector, index storage, and document refresh sit on separate line items, or a promotional first-term rate converts at renewal. Rebuild the total yourself from the line items, matching seat counts, contract length, and the exit position: what happens to the index and to the coaching history if you stop paying. NTT DATA's research on "enhanced humans" describes potential not limited by time, task, or knowledge; that promise does not remove the need to verify the specific thing you are signing.
| Mistake | What you would see | Pre-commit check |
|---|---|---|
| Testing the demo, not the live option | Clean citations from an admin account and a curated corpus | Run your question set on the production index as a standard employee |
| Counting seats as literacy | Enrollment numbers presented as trained staff | Request the completion report and the default exclusions list |
| Comparing unlike totals | Coaching fee quoted without retrieval and index costs | Rebuild the total from line items, terms, and exit position |
Write each verified result into the approval record, so the sign-off reflects what you tested rather than what you were shown.
Insider Tactics
The most useful insider move is to put the literacy material inside the retrieval corpus itself. Instead of a separate slide deck that ages out, publish your Article 4 guidance as a governed document in the same index the coaching tool searches, with a named owner and a review date. Then the coach answers "what may I paste into you?" from the source of truth. Check it directly: ask the live coach a policy question and confirm it cites your internal document rather than improvising a plausible-sounding answer. If it cannot, your people are learning how to use the tool somewhere other than the tool.
Pair that with point-of-use evidence. Your retrieval layer can log which documents informed each coaching answer and which user triggered it. Suprmind's guide to validated augmentation frames the goal cleanly: you set objectives, define quality standards, and approve outputs. Build that approval step into the coaching flow, and the same log does double duty as a retrieval-quality review and a dated record that a specific user received specific guidance. Align retention with your existing HR records policy instead of inventing a new one.
Scope literacy to permissions, not headcount. Retrieval is bounded by what a user's role can reach, so two people using the same coach may be drawing on different document sets. Write the literacy check for your narrowest and broadest roles, then sample a real query under each and confirm the retrieved set matches the training that role received. Uniform annual training misses this entirely; role-based checks catch it.
On timing, treat 2 February 2025 as history. Article 4's duty has applied since then, so in 2026 your leverage is change events, not the original date. Re-run the literacy check whenever the index, the embedding model, or the retrieval settings change, because a corpus swap can alter what the coach tells a user without anyone editing the training. Ask vendors to flag retrieval-affecting releases rather than every release.
Anchor refreshes to cycles you already run: procurement renewals and any performance cycle where coaching output could inform a review. Make the vendor's evidence pack a renewal condition rather than an afterthought, and timestamp your own pilot cohort so you hold first-party evidence if an auditor asks.
| Trigger | Tactic |
|---|---|
| Index, model, or retrieval settings change | Re-ask the policy question; confirm cited sources still match |
| Role or permission change | Re-check that the user's reachable corpus matches their literacy content |
| Contract renewal | Make the evidence pack a condition; refresh the literacy document |
Comparison
For a 2026 Article 4 review, compare two workable paths side by side: a vendor package built around a live retrieval-augmented workplace coach, and an internally delivered coaching program using the organization’s existing materials. The winner is the option that produces the complete, verified total for the live arrangement while giving relevant users usable evidence of instruction—not the option with the most attractive headline claim.
| Option | Known figure | Verification before commitment | When it wins |
|---|---|---|---|
| Vendor package with an organizational memory layer | Promoted as “9x cheaper” | Confirm the seller, live product version, included users, data limits, coaching scope, renewal terms, and all charges in the written offer. | It wins when the claimed saving applies to the exact workplace deployment and the package supplies complete, usable coaching evidence. |
| Internal coaching program | No grounded price available | Calculate staff time, preparation, delivery, maintenance, support, and any software or storage charges under the same user and coverage assumptions. | It wins when internal delivery can document the required instruction more completely than the vendor package at a lower verified total. |
The “9x cheaper” figure is not a comparable total by itself. Citybiz reported that Coworker.ai launched an organizational memory layer promising to make enterprise AI 9x cheaper, but that report does not establish the price of a particular coaching deployment. Treat the figure as a lead for diligence: request the baseline being compared, the billing period, and the exact included services before placing it beside an internal total.
For the vendor option, inspect the version that users will actually access after the February 2, 2025 AI-literacy duty took effect. Ask for a dated quote and a feature or service list that matches the proposed retrieval sources, user population, administrative controls, and coaching records. AWS describes RAG as a way businesses use external information with generative AI; that description supports checking the connected information sources, but it does not prove that a particular package includes the coaching or controls your review requires.
For the internal option, use one worksheet with identical assumptions: the same number of users, the same operating period, the same source-material scope, and the same support obligation. Record every included and excluded item, then compare the resulting totals rather than comparing a vendor marketing multiple with an unpriced internal estimate. If either side cannot provide a complete live total, the defensible choice is to delay commitment and obtain the missing terms.
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | Anchor the retrieval-coaching policy to the Article 4 AI-literacy duty start date, and record that date as the compliance baseline for workplace RAG use. | Article 4’s duty applies from that start date, so training and verification after it must be traceable to the same compliance anchor. |
| 2 | Apply the 60/40 train-or-verify split to retrieval coaching: use the 60 share to train users on RAG outputs, and the 40 share to verify retrieved outputs before commitment. | The split keeps coaching focused on both user capability and the live-output check that precedes any commitment. |
| 3 | Before committing to a RAG-assisted answer, require users to open the retrieved external source and confirm it is the live, complete option, not an incomplete or outdated version. | RAG incorporates external information into LLM responses, so a stale or partial retrieved source can make the final answer unreliable. |
| 4 | When the retrieved source presents totals or terms, compare like-for-like totals and terms across the same fields and scope before accepting the output. | This follows the decision rule: verify the live, complete option and compare like-for-like totals and terms; do not rely on an incomplete or outdated option. |
| 5 | In the 40 verification lane, sample completed RAG-assisted decisions and re-check the live, complete option before commitment; log any mismatch between the retrieved source and the output. | Sampling catches verification gaps and shows whether the 60/40 split is producing trustworthy retrieval coaching. |
| 6 | Have the 60 training lane document the verification habit: source retrieved, date checked, and whether the live, complete option was used before commitment. | Documentation turns the Article 4 AI-literacy duty into an auditable routine rather than a one-off instruction. |
Frequently Asked Questions
What does the train-or-verify approach require users to do with retrieved outputs?
It frames coaching around training users and verifying retrieved outputs before commitment.
What must be checked before committing to a retrieved option?
Users must verify that the option is live and complete before committing.
How should competing options be compared?
They should be compared using like-for-like totals and terms.
Which options should users avoid relying on?
Users should not rely on an incomplete or outdated option.
What does retrieval-augmented generation add to an LLM response?
RAG adds external information to LLM responses.
How does the RAG process use documents and model training?
The model first refers to a specified set of documents and then responds, supplementing what it learned in training.
Quick answers
| What does the guide treat as the compliance starting point for workplace AI literacy? | The guide treats that date as the compliance starting point for workplace AI literacy under the EU AI Act. |
| What does RAG add to LLM responses? | RAG adds external information to LLM responses. |
| How does retrieval-augmented generation work? | Retrieval-augmented generation retrieves relevant information from external data sources and incorporates it into responses. |
| What should users verify before committing? | Verify the live, complete option before committing. |
| How should options be compared? | Compare like-for-like totals and terms, and do not rely on an incomplete or outdated option. |
Also worth reading: Team coaching tools: Docebo vs Zoom video at 12-seat threshold: Team coaching tools: Docebo vs · New hire coaching: Retrieval-Augmented Coaching (RAC) cuts 31% vs mentor 2026: New hire coaching: Retrieval-Augmented Coaching · Employee training coaching 2026: Retrieval-Augmented Coaching (RAC) 28% vs mentor: Employee training coaching 2026: Retrieval-Augmented