05 — Zero-shot, One-shot & Few-shot Prompting¶
Learn how Zero-shot, One-shot, and Few-shot prompting enable Large Language Models (LLMs) to perform tasks with no examples, a single example, or multiple demonstrations — and understand how to select, design, evaluate, and implement these patterns in production AI systems.
📖 Overview¶
Large Language Models can perform many tasks without task-specific training.
However, the way instructions and examples are provided can significantly influence the model's behavior.
Three foundational prompting strategies are:
The progression can be visualized as:
flowchart LR
A["Zero-shot<br/>Instruction Only"] --> B["One-shot<br/>Instruction + 1 Example"]
B --> C["Few-shot<br/>Instruction + Multiple Examples"]
A --> D["LLM"]
B --> D
C --> D
These techniques are especially useful for:
- Classification
- Information extraction
- Text transformation
- Sentiment analysis
- Intent detection
- Formatting
- Content generation
- Domain-specific terminology
- Structured outputs
The important engineering question is not:
"Should I always use few-shot prompting?"
Instead:
Which prompting strategy provides sufficient quality for the task at an acceptable cost, latency, and complexity?
1. What Is Zero-shot Prompting?¶
Zero-shot prompting means asking an LLM to perform a task without providing task-specific examples.
The model receives:
and generates the result.
Example:
Classify the following customer message as
positive, negative, or neutral.
Message:
"The payment was successful."
There is no example showing how previous messages were classified.
2. Zero-shot Prompting Architecture¶
flowchart LR
A["Task Instruction"] --> C["Prompt"]
B["Input"] --> C
C --> D["LLM"]
D --> E["Output"]
The model relies on:
- Its pretrained knowledge
- The task description
- The supplied input
- The requested output format
3. Zero-shot Example — Sentiment Classification¶
Classify the following sentence as:
- positive
- negative
- neutral
Sentence:
"The application performed extremely well."
Possible output:
No examples were supplied.
4. Zero-shot Example — Intent Classification¶
Classify the following customer request into one of:
- PAYMENT
- ACCOUNT
- SHIPPING
- TECHNICAL
- OTHER
Customer request:
"My payment was deducted but the order was not created."
Possible output:
Again, the model receives only the instruction and input.
5. Zero-shot Example — Text Transformation¶
Convert the following sentence into a professional
business communication:
"Send me the report quickly."
Input:
"Send me the report quickly."
Possible output:
6. When Zero-shot Prompting Works Well¶
Zero-shot prompting is often effective when:
- The task is simple.
- The task is clearly defined.
- Categories have obvious meanings.
- The desired output is straightforward.
- The model already understands the task concept.
Examples:
7. Advantages of Zero-shot Prompting¶
Zero-shot prompting provides several advantages.
Simplicity¶
Low Token Usage¶
No demonstration examples are included.
Lower Latency¶
The prompt is generally smaller.
Lower Cost¶
Fewer input tokens may reduce inference cost.
Easier Maintenance¶
There are no example datasets embedded in the prompt.
8. Limitations of Zero-shot Prompting¶
Zero-shot prompting can become less reliable when:
- The task is ambiguous.
- The output format is unusual.
- Domain terminology is specialized.
- Categories are difficult to distinguish.
- The expected behavior is not obvious.
- The task contains implicit business rules.
Example:
The model does not know:
A more explicit prompt may solve the problem.
9. Zero-shot Prompting with Constraints¶
Zero-shot prompting does not mean the prompt must be minimal.
You can provide detailed instructions without providing examples.
Example:
Classify the customer message.
Allowed categories:
- PAYMENT
- ACCOUNT
- SHIPPING
- TECHNICAL
- OTHER
Rules:
- PAYMENT refers to transaction-related problems.
- ACCOUNT refers to login or account-management problems.
- SHIPPING refers to delivery-related problems.
- TECHNICAL refers to application or infrastructure issues.
- OTHER applies when none of the above categories match.
Return only the category name.
Message:
{{message}}
This is still zero-shot because:
10. Zero-shot vs No Instructions¶
Zero-shot does not mean providing no instructions.
Compare:
Poor¶
Zero-shot¶
Classify the following message as
PAYMENT, ACCOUNT, SHIPPING, TECHNICAL, or OTHER.
Message:
"The payment failed."
The second prompt is zero-shot.
It simply does not provide demonstrations.
11. What Is One-shot Prompting?¶
One-shot prompting provides one example of the desired task behavior.
The structure becomes:
Example:
Classify the sentiment.
Example:
Input:
"The service was excellent."
Output:
positive
Now classify:
Input:
"The application keeps crashing."
Possible output:
12. One-shot Architecture¶
flowchart TD
A["Task Instruction"] --> D["Prompt"]
B["Example Input + Output"] --> D
C["New Input"] --> D
D --> E["LLM"]
E --> F["Output"]
The example acts as a demonstration of the expected behavior.
13. Why Use One-shot Prompting?¶
One example can clarify:
For example, suppose the task is:
The meaning of severity may not be obvious.
A demonstration can clarify the intended classification.
14. One-shot Example — Incident Classification¶
Classify the incident as:
- LOW
- MEDIUM
- HIGH
Example:
Input:
"One user experienced a temporary UI error."
Output:
LOW
Now classify:
Input:
"All payment transactions are failing."
Possible output:
The example provides a reference point.
15. One-shot Example — Structured Output¶
Extract the customer information.
Example:
Input:
"John placed order ORD-1001."
Output:
{
"customer_name": "John",
"order_id": "ORD-1001"
}
Now process:
Input:
"Sarah placed order ORD-1002."
Expected:
The example demonstrates both:
16. One-shot Example — Transformation¶
Convert technical language into language
appropriate for a business stakeholder.
Example:
Input:
"The service experienced a database connection pool exhaustion."
Output:
"The application temporarily could not connect
to the database because available connections were exhausted."
Now convert:
Input:
"The Kafka consumer lag increased significantly."
The model now has one demonstration of the desired transformation style.
17. Advantages of One-shot Prompting¶
One-shot prompting can provide:
- Better task clarification
- Better formatting guidance
- Better style consistency
- More predictable output
- Lower prompt complexity than many-shot approaches
It can be a useful middle ground between:
and:
18. Limitations of One-shot Prompting¶
A single example may not capture:
For example:
does not necessarily teach the model how to distinguish:
If the task is complex, multiple examples may be more useful.
19. What Is Few-shot Prompting?¶
Few-shot prompting provides multiple examples before the new input.
The structure becomes:
Example:
Classify sentiment.
Example 1:
Input:
"The service was excellent."
Output:
positive
Example 2:
Input:
"The application was unavailable."
Output:
negative
Example 3:
Input:
"The application is running normally."
Output:
neutral
Now classify:
Input:
"The response time has improved significantly."
Possible output:
20. Few-shot Architecture¶
flowchart TD
A["Task Instruction"] --> F["Prompt"]
B["Example 1"] --> F
C["Example 2"] --> F
D["Example 3"] --> F
E["Example N"] --> F
G["New Input"] --> F
F --> H["LLM"]
H --> I["Output"]
The examples establish a small set of demonstrations for the model.
21. Why Few-shot Prompting Works¶
Few-shot examples can communicate patterns that are difficult to describe using instructions alone.
For example:
The model can infer the intended mapping.
This is particularly useful when:
or:
22. Few-shot Classification Example¶
Classify each customer message.
Categories:
- PAYMENT
- ACCOUNT
- SHIPPING
- TECHNICAL
Examples:
Input:
"My credit card was charged twice."
Output:
PAYMENT
Input:
"I cannot reset my password."
Output:
ACCOUNT
Input:
"My package has not arrived."
Output:
SHIPPING
Input:
"The application crashes when I upload a file."
Output:
TECHNICAL
Now classify:
Input:
"My card was charged but the order failed."
Possible output:
23. Few-shot Extraction Example¶
Extract:
- customer_name
- order_id
- amount
Example 1:
Input:
"John ordered item ORD-1001 for $200."
Output:
{
"customer_name": "John",
"order_id": "ORD-1001",
"amount": 200
}
Example 2:
Input:
"Sarah purchased order ORD-1002 for $350."
Output:
{
"customer_name": "Sarah",
"order_id": "ORD-1002",
"amount": 350
}
Now process:
Input:
"Michael purchased order ORD-1003 for $175."
Possible output:
24. Few-shot Formatting Example¶
Few-shot prompting is particularly useful for custom formatting.
Convert the input into the requested format.
Example 1:
Input:
"Java backend developer"
Output:
{
"role": "Backend Developer",
"primary_language": "Java"
}
Example 2:
Input:
"Python data scientist"
Output:
{
"role": "Data Scientist",
"primary_language": "Python"
}
Now convert:
Input:
"AWS cloud architect"
Expected:
The examples communicate the desired representation.
25. Zero-shot vs One-shot vs Few-shot¶
The core difference is the number of demonstrations.
| Strategy | Examples | Complexity | Token Usage |
|---|---|---|---|
| Zero-shot | 0 | Low | Low |
| One-shot | 1 | Medium | Medium |
| Few-shot | 2+ | Higher | Higher |
Conceptually:
flowchart LR
A["Zero-shot<br/>0 Examples"] --> B["One-shot<br/>1 Example"]
B --> C["Few-shot<br/>Multiple Examples"]
A --> D["Lowest Prompt Overhead"]
B --> E["Moderate Prompt Overhead"]
C --> F["Highest Prompt Overhead"]
26. Choosing the Right Strategy¶
A practical decision process:
flowchart TD
A["Define Task"] --> B{"Is Task Simple?"}
B -->|Yes| C["Try Zero-shot"]
B -->|No| D["Try One-shot"]
C --> E{"Quality Acceptable?"}
D --> F{"Quality Acceptable?"}
E -->|Yes| G["Use Zero-shot"]
E -->|No| H["Add One or More Examples"]
F -->|Yes| I["Use One-shot"]
F -->|No| H
H --> J["Few-shot"]
J --> K["Evaluate"]
A useful engineering principle is:
Start with the simplest prompting strategy that meets the quality requirement.
27. Prompt Escalation Strategy¶
A production workflow can use:
This avoids unnecessarily adding examples and token cost.
28. Example of Prompt Escalation¶
Version 1 — Zero-shot¶
If results are inconsistent:
Version 2 — One-shot¶
Classify the incident as LOW, MEDIUM, or HIGH.
Example:
Input:
"One user experienced a temporary UI issue."
Output:
LOW
Incident:
{{incident}}
If quality is still insufficient:
Version 3 — Few-shot¶
Example 1:
...
Output: LOW
Example 2:
...
Output: MEDIUM
Example 3:
...
Output: HIGH
Incident:
{{incident}}
29. Example Selection¶
Few-shot prompting is not simply:
Example selection is critical.
Good examples should be:
- Relevant
- Representative
- Correct
- Diverse
- Clear
- Consistent
Poor examples can teach the wrong behavior.
30. Representative Examples¶
Suppose categories are:
A weak few-shot set might contain:
A better set contains:
This gives the model examples across the task space.
31. Diverse Examples¶
Examples should cover meaningful variations.
For payment classification:
This is more useful than five nearly identical examples.
32. Boundary Examples¶
The most difficult cases can be especially valuable.
Example:
This may involve both:
A boundary example can clarify the expected category.
33. Example Quality¶
Every demonstration should be verified.
Bad:
If the example is incorrect, the model receives misleading supervision.
Therefore:
Few-shot examples are part of the prompt and should be treated as production artifacts.
34. Example Ordering¶
The ordering of demonstrations can influence model behavior.
A simple structure is:
Keep the structure consistent.
If examples have inconsistent formatting:
the model receives conflicting output signals.
35. Consistent Demonstration Format¶
Prefer:
for every example.
Example:
Example 1
Input:
Payment failed.
Output:
PAYMENT
Example 2
Input:
Password reset failed.
Output:
ACCOUNT
Then:
36. Positive and Negative Examples¶
Examples can demonstrate both successful and problematic cases.
For example:
and:
Together they demonstrate category boundaries.
37. Few-shot Examples and Output Formatting¶
Examples can be used to demonstrate exact output structure.
For example:
This can be more effective than simply saying:
because the model sees the exact desired representation.
38. Few-shot and Structured Output¶
A production pattern is:
Architecture:
flowchart LR
A["Few-shot Examples"] --> D["Prompt"]
B["New Input"] --> D
C["Output Schema"] --> D
D --> E["LLM"]
E --> F["Validator"]
F --> G["Application"]
The examples guide behavior.
The schema and validator provide stronger application-level control.
39. Few-shot Prompt Template¶
A reusable template can be:
SYSTEM
You are an enterprise classification assistant.
TASK
Classify the input.
CATEGORIES
{{categories}}
EXAMPLES
{{examples}}
INPUT
{{input}}
OUTPUT
Return only the category.
The runtime application can inject:
40. Python Few-shot Template¶
examples = """
Example 1:
Input:
"My card was charged twice."
Output:
PAYMENT
Example 2:
Input:
"I cannot reset my password."
Output:
ACCOUNT
Example 3:
Input:
"My package has not arrived."
Output:
SHIPPING
"""
prompt = f"""
You are an enterprise customer-support classifier.
Classify the following message.
Categories:
PAYMENT
ACCOUNT
SHIPPING
TECHNICAL
OTHER
Examples:
{examples}
Input:
{user_message}
Return only the category.
"""
The application dynamically constructs the prompt.
41. Framework Example — LangChain¶
A few-shot prompt can be implemented using LangChain prompt abstractions.
from langchain_core.prompts import (
ChatPromptTemplate,
FewShotChatMessagePromptTemplate
)
example_prompt = ChatPromptTemplate.from_messages([
("human", "{input}"),
("ai", "{output}")
])
examples = [
{
"input": "My card was charged twice.",
"output": "PAYMENT"
},
{
"input": "I cannot reset my password.",
"output": "ACCOUNT"
},
{
"input": "My package has not arrived.",
"output": "SHIPPING"
}
]
few_shot_prompt = FewShotChatMessagePromptTemplate(
example_prompt=example_prompt,
examples=examples
)
prompt = ChatPromptTemplate.from_messages([
(
"system",
"""
You are an enterprise customer-support classifier.
Return only:
PAYMENT, ACCOUNT, SHIPPING, TECHNICAL, or OTHER.
"""
),
few_shot_prompt,
("human", "{input}")
])
messages = prompt.invoke({
"input": "My payment was deducted but the order failed."
})
The framework provides prompt composition.
The underlying concept remains:
Detailed LangChain architecture is intentionally covered later in Part VIII — AI Engineering Frameworks & Tooling.
42. Framework Example — LlamaIndex¶
A simple LlamaIndex-style prompt can also be constructed with examples.
from llama_index.core import PromptTemplate
template = PromptTemplate(
"""
You are an enterprise classifier.
Examples:
Example 1:
Input: My card was charged twice.
Output: PAYMENT
Example 2:
Input: I cannot reset my password.
Output: ACCOUNT
Classify:
Input:
{input}
Return only the category.
"""
)
prompt = template.format(
input="My password reset link has expired."
)
Again, the core pattern is framework-independent.
43. Dynamic Few-shot Examples¶
Production applications may select examples dynamically.
Instead of:
the system can:
Conceptually:
flowchart TD
A["User Query"] --> B["Example Selection"]
C["Example Store"] --> B
B --> D["Selected Examples"]
D --> E["Prompt Builder"]
A --> E
E --> F["LLM"]
This can reduce irrelevant examples while preserving few-shot guidance.
44. Similarity-Based Example Selection¶
Suppose the application has:
Sending all examples is impractical.
Instead:
This is conceptually similar to retrieval.
Example:
Query:
"Card charged but payment status failed."
Selected examples:
1. Card charged twice → PAYMENT
2. Card charged but order failed → PAYMENT
3. Refund not received → PAYMENT
This approach is sometimes called dynamic few-shot prompting.
45. Few-shot Prompting vs RAG¶
Few-shot prompting and RAG solve different problems.
Few-shot prompting¶
Provides:
RAG¶
Provides:
Conceptually:
flowchart TD
A["LLM Application"] --> B["Behavior Guidance"]
A --> C["Knowledge Retrieval"]
B --> D["Few-shot Examples"]
C --> E["Retrieved Context"]
D --> F["Prompt"]
E --> F
F --> G["LLM"]
They can also be combined.
46. Few-shot + RAG¶
A RAG application may use:
For example:
SYSTEM
You are an enterprise knowledge assistant.
EXAMPLES
Example 1:
...
Example 2:
...
CONTEXT
{{retrieved_context}}
QUESTION
{{question}}
OUTPUT
Answer using only the supplied context.
This combines:
47. Few-shot Prompting and Token Cost¶
Every example consumes input tokens.
If:
and:
then the prompt becomes substantially larger.
At scale:
flowchart LR
A["Number of Examples"] --> B["Prompt Tokens"]
B --> C["Request Cost"]
B --> D["Latency"]
Therefore:
More examples are not automatically better.
48. Example Count Optimization¶
Suppose evaluation produces:
| Examples | Accuracy | Tokens |
|---|---|---|
| 0 | 82% | 300 |
| 1 | 87% | 400 |
| 3 | 91% | 600 |
| 5 | 92% | 800 |
| 10 | 92% | 1,300 |
The difference between:
and:
may not justify the additional token cost.
The actual decision should be based on measured application behavior.
49. Example Selection Trade-offs¶
Good few-shot selection balances:
A useful conceptual model:
flowchart TD
A["Candidate Examples"] --> B["Relevance"]
A --> C["Coverage"]
A --> D["Diversity"]
B --> E["Example Selector"]
C --> E
D --> E
E --> F["Token Budget"]
F --> G["Selected Examples"]
50. Few-shot Prompting and Context Windows¶
Few-shot examples consume part of the model's context window.
A simplified prompt budget is:
System Instructions
+
Few-shot Examples
+
Retrieved Context
+
Conversation History
+
User Input
+
Output
All of these compete for available context.
Therefore:
Few-shot prompting must be designed together with context management.
51. Few-shot and Long Context¶
A common mistake is:
This can produce an unnecessarily large prompt.
A better approach is:
The objective is to maximize useful information rather than raw context size.
52. Example Selection Rules¶
A production example selector can use rules such as:
1. Select examples from the same task.
2. Prefer examples similar to the current input.
3. Cover different categories.
4. Avoid contradictory demonstrations.
5. Remove redundant examples.
6. Respect a token budget.
7. Validate example correctness.
53. Few-shot Example Versioning¶
Examples should be versioned alongside prompts.
Example:
prompts/
└── customer-classification/
├── v1/
│ ├── prompt.txt
│ └── examples.json
│
└── v2/
├── prompt.txt
└── examples.json
This allows production behavior to be reproduced.
54. Example Dataset¶
Examples can be stored externally.
[
{
"input": "My card was charged twice.",
"output": "PAYMENT"
},
{
"input": "I cannot reset my password.",
"output": "ACCOUNT"
},
{
"input": "My package has not arrived.",
"output": "SHIPPING"
}
]
The application can load the examples at runtime.
This is preferable to embedding a very large example collection directly into source code.
55. Few-shot Prompt Builder¶
def build_few_shot_prompt(
examples,
user_input
):
formatted_examples = []
for example in examples:
formatted_examples.append(
f"""
Input:
{example["input"]}
Output:
{example["output"]}
"""
)
examples_text = "\n".join(formatted_examples)
return f"""
You are an enterprise classifier.
Examples:
{examples_text}
Input:
{user_input}
Return only the classification.
"""
This separates:
from:
56. Production Architecture¶
A production few-shot system may look like:
flowchart TD
A["User Request"] --> B["Application API"]
B --> C["Task Classifier"]
C --> D["Example Selector"]
E["Example Store"] --> D
D --> F["Selected Examples"]
F --> G["Prompt Builder"]
B --> G
G --> H["LLM Provider"]
H --> I["Output Parser"]
I --> J["Validation"]
J --> K["Application Response"]
G --> L["Prompt Observability"]
H --> L
J --> L
This architecture separates:
57. Zero-shot, One-shot, Few-shot — Production Comparison¶
| Dimension | Zero-shot | One-shot | Few-shot |
|---|---|---|---|
| Examples | None | One | Multiple |
| Prompt Size | Smallest | Small | Larger |
| Implementation | Simplest | Simple | More involved |
| Cost | Lowest | Low | Higher |
| Latency | Lower | Low | Higher |
| Task Guidance | Low | Medium | High |
| Format Guidance | Limited | Better | Strong |
| Domain Adaptation | Limited | Moderate | Better |
| Maintenance | Easy | Easy | Higher |
| Example Management | None | One | Required |
These are general architectural tendencies. Actual performance depends on the model, task, prompt, examples, and deployment environment.
58. When to Use Zero-shot¶
Prefer zero-shot when:
Examples:
59. When to Use One-shot¶
Consider one-shot when:
Examples:
60. When to Use Few-shot¶
Consider few-shot when:
Zero-shot quality is insufficient
+
One example is insufficient
+
Examples can demonstrate the required behavior.
Examples:
Complex classification
Domain-specific formatting
Subtle intent detection
Custom extraction rules
Specialized terminology
61. When Few-shot Is Not the Right Solution¶
Few-shot prompting is not a universal solution.
If the problem is:
Adding examples may not solve it.
The better solution may be:
If the problem is:
use:
If the problem is:
consider:
62. Prompt Strategy Decision Matrix¶
| Problem | Recommended First Approach |
|---|---|
| Simple task | Zero-shot |
| Ambiguous task | Better instructions |
| Output format unclear | One-shot example |
| Subtle classification | Few-shot |
| Domain knowledge missing | RAG / retrieval |
| Invalid structure | Structured output + validation |
| Complex workflow | Task decomposition |
| Large knowledge base | Retrieval |
| Repeated specialized task | Few-shot + evaluation |
| Security-sensitive action | Application-level controls |
63. Evaluation Workflow¶
Zero-shot, one-shot, and few-shot strategies should be compared experimentally.
flowchart TD
A["Evaluation Dataset"] --> B["Zero-shot"]
A --> C["One-shot"]
A --> D["Few-shot"]
B --> E["Metrics"]
C --> E
D --> E
E --> F["Compare Quality"]
F --> G["Compare Cost"]
G --> H["Compare Latency"]
H --> I["Production Decision"]
64. Evaluation Metrics¶
Depending on the task, evaluate:
Accuracy
Precision
Recall
F1
Exact Match
Format Compliance
Groundedness
Human Evaluation
Latency
Token Usage
Cost
For classification:
For structured extraction:
For generation:
65. Example Evaluation Dataset¶
test_cases = [
{
"input": "My card was charged twice.",
"expected": "PAYMENT"
},
{
"input": "I cannot reset my password.",
"expected": "ACCOUNT"
},
{
"input": "My package has not arrived.",
"expected": "SHIPPING"
},
{
"input": "The application crashes on startup.",
"expected": "TECHNICAL"
}
]
Run the same dataset against:
Then compare the results.
66. Prompt Regression Testing¶
Once a few-shot prompt is deployed, changing the examples can change behavior.
Therefore:
flowchart LR
A["Prompt + Examples v1"] --> C["Regression Dataset"]
B["Prompt + Examples v2"] --> C
C --> D["Compare Metrics"]
D --> E["Release Decision"]
Examples are part of the prompt behavior and must be included in regression testing.
67. Prompt Versioning¶
A production prompt registry might contain:
customer-classification/
├── v1/
│ ├── prompt.md
│ └── examples.json
│
├── v2/
│ ├── prompt.md
│ └── examples.json
│
└── v3/
├── prompt.md
└── examples.json
A production trace should identify:
68. Observability¶
For production applications, track:
Prompt Version
Example Set
Model
Input Tokens
Output Tokens
Latency
Validation Result
Retry Count
Final Result
Conceptually:
flowchart TD
A["Request"] --> B["Prompt"]
B --> C["Example Set"]
B --> D["Model"]
C --> E["Observability"]
D --> E
B --> E
F["Token Usage"] --> E
G["Latency"] --> E
H["Validation"] --> E
69. Cost Optimization¶
Few-shot examples increase input token consumption.
Optimization strategies include:
Use fewer examples
+
Remove redundant examples
+
Select relevant examples
+
Compress examples
+
Use zero-shot when sufficient
A practical optimization loop is:
70. Latency Optimization¶
If a few-shot prompt contains many examples:
Therefore:
are preferable to:
71. Security Considerations¶
Few-shot examples are still prompt content.
If examples are dynamically loaded from external systems, they should be treated carefully.
Potential sources include:
A malicious example could attempt to influence model behavior.
Therefore:
flowchart TD
A["Example Source"] --> B["Validation"]
B --> C["Example Selection"]
C --> D["Prompt Builder"]
D --> E["LLM"]
Examples should be trusted, validated, and controlled.
72. Few-shot Prompt Injection Risk¶
Suppose an example contains:
If the example is inserted into a production prompt without validation, it may create unintended behavior.
Therefore:
Few-shot examples should be treated as controlled prompt assets, not arbitrary text.
73. Example Data Governance¶
Enterprise applications should consider:
Do not blindly use historical customer conversations as examples.
Examples may contain:
Example datasets should be appropriately sanitized and governed.
74. Production Workflow¶
A production few-shot workflow can be:
1. Define the task.
2. Try zero-shot prompting.
3. Create evaluation dataset.
4. Measure baseline quality.
5. Add one representative example.
6. Evaluate again.
7. Add more examples only if necessary.
8. Select relevant and diverse examples.
9. Validate all examples.
10. Define a token budget.
11. Version the prompt and example set.
12. Deploy.
13. Monitor quality, latency, and cost.
14. Run regression tests after changes.
75. Practical Example — Customer Support¶
Zero-shot¶
One-shot¶
Classify the customer request.
Example:
Input:
"I cannot reset my password."
Output:
ACCOUNT
Request:
{{request}}
Few-shot¶
Classify the customer request.
Example 1:
Input:
"My card was charged twice."
Output:
PAYMENT
Example 2:
Input:
"I cannot reset my password."
Output:
ACCOUNT
Example 3:
Input:
"My package has not arrived."
Output:
SHIPPING
Example 4:
Input:
"The application crashes when I upload a file."
Output:
TECHNICAL
Request:
{{request}}
The application can evaluate all three versions and choose the simplest one that satisfies the quality requirement.
76. Practical Example — Document Extraction¶
Zero-shot¶
One-shot¶
Extract customer_name, order_id, and amount.
Example:
Input:
"John placed order ORD-1001 for $200."
Output:
{
"customer_name": "John",
"order_id": "ORD-1001",
"amount": 200
}
Document:
{{document}}
Few-shot¶
Extract customer_name, order_id, and amount.
Example 1:
Input:
"John placed order ORD-1001 for $200."
Output:
{
"customer_name": "John",
"order_id": "ORD-1001",
"amount": 200
}
Example 2:
Input:
"Sarah purchased ORD-1002 for $350."
Output:
{
"customer_name": "Sarah",
"order_id": "ORD-1002",
"amount": 350
}
Document:
{{document}}
77. Practical Example — Backend Engineering¶
Suppose an enterprise AI assistant needs to classify backend incidents.
Zero-shot¶
One-shot¶
Example:
Incident:
"Consumer lag increased across all Kafka partitions."
Category:
KAFKA
Incident:
{{incident}}
Few-shot¶
Example 1:
Incident:
"Consumer lag increased across all Kafka partitions."
Category:
KAFKA
Example 2:
Incident:
"PostgreSQL connections are exhausted."
Category:
DATABASE
Example 3:
Incident:
"API requests return HTTP 503."
Category:
API
Example 4:
Incident:
"Unauthorized access attempts were detected."
Category:
SECURITY
Incident:
{{incident}}
The examples establish the mapping between symptoms and categories.
78. Prompt Pattern with Few-shot Examples¶
A reusable production template:
ROLE
You are an enterprise AI assistant.
TASK
{{task}}
RULES
{{rules}}
EXAMPLES
{{examples}}
INPUT
{{input}}
OUTPUT
{{output_format}}
This can support:
79. Zero-shot, One-shot, Few-shot — Decision Summary¶
Start
↓
Try Zero-shot
↓
Evaluate
↓
Quality sufficient?
┌───────────────┐
│ Yes │ No
↓ ↓
Deploy Try One-shot
↓
Evaluate
↓
Quality sufficient?
┌─────────────┐
│ Yes │ No
↓ ↓
Deploy Try Few-shot
↓
Evaluate
↓
Deploy
This strategy keeps prompt complexity proportional to actual requirements.
80. Common Mistakes¶
80.1 Adding Examples Without Measuring Them¶
More examples do not automatically improve quality.
80.2 Using Irrelevant Examples¶
An example should be related to the task.
80.3 Using Incorrect Examples¶
Incorrect demonstrations can reinforce incorrect behavior.
80.4 Using Too Many Examples¶
Excessive examples increase:
80.5 Redundant Examples¶
Ten examples that demonstrate exactly the same case may add little value.
80.6 Inconsistent Formatting¶
Examples should follow a consistent structure.
80.7 Ignoring Edge Cases¶
Examples should cover important boundary conditions.
80.8 Using Few-shot Instead of Retrieval¶
Examples teach behavior.
They do not replace current enterprise knowledge.
80.9 Treating Few-shot Examples as Security Controls¶
Examples should never replace application security.
81. Best Practices¶
1. Start with zero-shot.
2. Measure the baseline.
3. Add one example when behavior or format needs clarification.
4. Use few-shot when multiple demonstrations materially improve performance.
5. Choose representative examples.
6. Prefer diverse examples.
7. Include boundary cases when useful.
8. Keep example formatting consistent.
9. Validate example correctness.
10. Remove redundant examples.
11. Respect the context/token budget.
12. Version prompts and examples.
13. Evaluate prompt changes using regression datasets.
14. Monitor quality, latency, and cost.
15. Sanitize sensitive example data.
16. Use retrieval when the problem is missing knowledge rather than missing behavioral guidance.
17. Validate final outputs at the application layer.
82. Key Takeaways¶
- Zero-shot prompting uses instructions without demonstrations.
- One-shot prompting uses one demonstration.
- Few-shot prompting uses multiple demonstrations.
- Zero-shot is usually the simplest and most token-efficient approach.
- One-shot can clarify unusual task behavior or output formats.
- Few-shot can improve performance on complex or specialized tasks.
- More examples do not automatically produce better results.
- Example quality is often more important than example quantity.
- Good examples should be:
- Relevant
- Representative
- Diverse
- Correct
- Consistently formatted
- Few-shot examples consume context-window capacity and tokens.
- Few-shot prompting should therefore be evaluated against cost and latency.
- Dynamic example selection can retrieve relevant demonstrations at runtime.
- Few-shot prompting and RAG solve different problems:
- Few-shot → demonstrates behavior
- RAG → supplies knowledge
- They can be combined in a single application.
- Prompt and example sets should be versioned together.
- Example data should be validated and governed.
- Sensitive enterprise data should not be blindly embedded into examples.
- Prompting should be evaluated systematically rather than based on one successful response.
- The practical strategy is:
The central production principle is:
Use the minimum number of high-quality examples necessary to achieve the required behavior.
83. Chapter Navigation¶
Part IV — Prompt Engineering & RAG Fundamentals¶
Previous Chapter: 04. Prompt Design Patterns
Current Chapter: 05 — Zero-shot, One-shot & Few-shot Prompting
Next Chapter: 06. Chain-of-Thought Prompting
Part IV Chapters¶
- 01. Introduction to Prompt Engineering
- 02. Prompt Engineering Fundamentals
- 03. Advanced Prompt Engineering
- 04. Prompt Design Patterns
- 05. Zero-shot, One-shot & Few-shot Prompting
- 06. Chain-of-Thought Prompting
- 07. ReAct Prompting
- 08. Structured Outputs & Output Parsing
- 09. Function Calling & Tool Calling
- 10. Embeddings in Practice
- 11. Document Processing & Vectorization
- 12. Document Chunking Strategies
- 13. Vector Database Fundamentals
- 14. Similarity Search Techniques
- 15. RAG Pipeline Components
- 16. Retrieval & Generation Pipeline
- 17. Vector Databases in RAG
- 18. Building Your First RAG Pipeline
- 19. RAG Evaluation Fundamentals
- 20. Enterprise Generative AI Application Architecture
- 21. Deploying AI Applications with Gradio
References¶
- OpenAI — Prompt Engineering and API Documentation
- Anthropic — Prompt Engineering Documentation
- Google — Gemini API and Generative AI Documentation
- Hugging Face — Transformers Documentation
- LangChain — Prompt Templates and Few-shot Prompting Documentation
- LlamaIndex — Prompt Templates and LLM Application Documentation
- Brown et al. — Language Models are Few-Shot Learners
- Wei et al. — Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Min et al. — Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
- Liu et al. — What Makes Good In-Context Examples for GPT-3?
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems — One Chapter at a Time.