Appendix Appendix A Full Annotation Prompts
This appendix reproduces the full prompts used in the two-stage annotation pipeline. Stage 1 uses Llama-3.1-8b-instant for task extraction and task–skill matching. Stage 2 uses GPT-5-nano for task-level exposure classification.
A.1 Stage 1 Prompt: Task Extraction and Task–Skill Matching
You are an expert in job task analysis and workplace activities.``Your goal is to extract REAL, POSTING-SPECIFIC tasks from individual job postings.``These tasks will later be used to evaluate generative-AI exposure, so accuracy and realism are critical.``Analyze each job posting and output EXACTLY ONE valid JSON object matching the schema below.``Return ONLY JSON — no markdown, no notes, no explanations.``Each posting provides:``- ID: unique identifier``- TITLE_NAME: job title``- BODY: job description text``- SPECIALIZED_SKILLS_NAME: list of technical / hard skills``- COMMON_SKILLS_NAME: list of general / soft skills``IMPORTANT GUIDANCE:``- You are NOT summarizing an occupation or a typical role.``- You are extracting what THIS specific job posting describes workers actually doing.``- Do NOT generalize, normalize, or rewrite tasks into abstract or generic duties.``- Do NOT invent tasks that are not grounded in the posting text.``- Tasks should be traceable to the BODY text (verbatim or lightly edited for clarity).``STEP 1. Group Skills (Supportive Context Only)``- Group skills by semantic similarity within their original type.``- Specialized skills → S1, S2, ...``- Common skills → C1, C2, ...``- Each group: {"group_id": "...", "group_skills": [up to 5 similar skills]}.``- Do NOT mix specialized and common skills in the same group.``- If no skills are provided, create exactly one group:``{"group_id": "NS0", "group_skills": []}.``STEP 2. Extract Posting-Specific Tasks (HIGH FIDELITY)``Extract 3–10 concrete, actionable tasks describing what the worker ACTUALLY DOES in this posting.``Task extraction rules:``- Tasks must be grounded in the BODY text and reflect real actions, tools, systems,``materials, interactions, or work settings mentioned.``- Preserve specificity whenever possible (software names, equipment, documents,``samples, customers, patients, systems, environments).``- Focus on observable actions (clear verbs + objects).``- Only include tasks explicitly stated or clearly implied by the BODY text.``STRICTLY AVOID:``- Company information, compensation, benefits, qualifications, or hiring instructions.``--------------------------------------------------``STEP 3. Match Tasks to Skill Groups``--------------------------------------------------``- Assign EACH task exactly ONE skill_group_id.``- Compare the task against all skill groups (specialized + common).``- Choose the closest semantic match.``- In case of ties, prefer specialized groups (S*) over common groups (C*).``- If no skills exist, assign all tasks to NS0.``--------------------------------------------------``OUTPUT FORMAT (STRICT — JSON ONLY)``--------------------------------------------------``Return EXACTLY one JSON object with this schema:``{``"posting_id": "<ID>",``"posting_title": "<TITLE_NAME>",``"skills_groups": [``{"group_id": "S1", "group_skills": ["..."]},``{"group_id": "C1", "group_skills": ["..."]},``{"group_id": "NS0", "group_skills": []}``],``"tasks": [``{``"task_id": "t1",``"task": "<8–50-word posting-specific task>",``"skill_group_id": "S1 | C1 | NS0"``}``]``}``FINAL CHECKS:``- Valid JSON only.``- 3–10 tasks.``- Tasks reflect concrete work described in the posting, not generalized role descriptions.``- Every task has exactly one skill_group_id.
A.2 Stage 2 Prompt: Exposure Classification
You are an expert on Generative AI. Your task is to classify work tasks by their exposure to GenAI/LLM tools.``Assume:``- A worker with average expertise in this role.``- Access to powerful LLMs and readily available GenAI tools.``- Access to standard laptop tools (e.g., microphone, speakers).``- NO physical tools or physical presence beyond a laptop.``You will receive ONE JSON object containing:``- posting_id: string. Job posting ID.``- posting_title: string. Job title.``- tasks: list of task objects, each with:``- task_id: string (e.g., "t1")``- task: string describing the worker task activity.``Assign exactly ONE exposure_label to each task in { "E0", "E1", "E2" }:``E0 (No exposure):``- Label E0 if readily available GenAI/LLM tools (ChatGPT, copilots, ASR/TTS, multimodal tools) cannot reduce task time by ≥50% at equivalent quality.``- Tasks requiring substantial in-person interaction, physical manipulation, inspections, repairs, hands-on care, or credential-bound decisions must be E0.``E1 (Direct exposure):``- Label E1 if a single off-the-shelf GenAI/LLM tool, with no special integrations or fine-tuning, can reduce effort by ≥50% at equivalent quality.``- Typical E1 patterns (not exhaustive):``* Writing or transforming text or code``* Translation, tone/style edits``* Summarizing medium-length documents; extracting structured info``* Drafting emails, briefs, slides, Q&A based on provided content``* ASR transcription; TTS draft narration; simple audio clean-up``* Simple image generation/editing (thumbnails, captions, marketing images)``- Not E1: tasks needing internal data access, enterprise systems, custom workflows, or automation beyond copy-paste.``E2 (LLM+ exposure):``- Label E2 if a single off-the-shelf GenAI/LLM tool alone cannot achieve ≥50% time reduction, but a thin AI-powered software layer could plausibly do so.``- Examples of thin layers:``* Retrieval or search over private/internal documents``* Multi-step workflows (draft → review → format → upload)``* In-product assist (CRM, IDE, helpdesk) that reads context and suggests actions/replies.``* Auto-suggest replies in customer service systems``* Rubric/policy/brand enforcement; light fine-tuning``* Integrated multimodal tools tied to asset libraries or translation pipelines``- Not E2:``* If E1 already applies → choose E1``* If even with integration AI cannot plausibly halve time → choose E0``* If heavy bespoke ML, robotics, or new model training is required → choose E0``OUTPUT FORMAT (STRICT):``Return ONE JSON object:``{``"posting_id": "<same as input>",``"task_exposures": [``{"task_id": "t1", "exposure_label": "E1"},``{"task_id": "t2", "exposure_label": "E0"}``]``}``Rules:``- Return ONLY valid JSON. No explanation or extra text.``- Every input task_id must appear exactly once.``- exposure_label must be exactly one of "E0", "E1", or "E2".