After using AI agents for data processing work for a while, I ran into a problem: the agent does a great job when the workload is small. But once I give it a bigger workload, it starts showing signs of trouble, like an overloaded context, trying to take shortcuts, and forgetting the rules I gave it. That pushed me to learn about subagents, and I came up with three strategies for using them to handle large workloads.
What is a subagent
The way I understand it, a subagent is a helper agent called by the main agent of your chat session. It works on the prompt the main agent gives it, then reports the result back so the main agent can put everything together and give you an overall summary. I currently use subagents in Claude Code, Anthropic’s AI agent tool.
A subagent can run on a different model from the main agent. For example, your chat session may use Opus 5.5 (10/2026) for the main agent, while the subagents run on Sonnet 5.5 to do the work. I like this setup, because tasks that don’t need much thinking can go to a lighter model to save tokens, while the main agent, which coordinates everything, uses a stronger model to reason and review the results.
A subagent can’t see your conversation, so when the main agent hands off a task, the prompt has to be complete. The subagent’s results also need to be checked and combined by the main agent before they are reported back to you.
3 strategies for using subagents
Subagents can work in parallel with the main agent of your chat session. But after applying them to a few projects, I found that the right strategy depends on how you split the data for processing and how the results are saved to storage.
Strategy 1 - Use subagents to process sequentially
Processing sequentially with subagents doesn’t make the work much faster, but it keeps the context window from getting overloaded, so each subagent sticks to the skills and rules more closely. So why sequential?
Let me share an example from a project I worked on. The main task was to classify messages in group chats to find out what requests customers made, what the customer service staff should do, and what the current status of each request is.
In a group chat, the agent has to read back a few earlier messages to understand the context of the message it is analyzing, and check the tickets to figure out which ticket the customer and the staff are replying to. That’s why I had the agent process messages sequentially in batches of 20 to 30 so its context wouldn’t get overloaded.
This task can’t run on parallel subagents, because the messages being processed depend on the results of the earlier messages. Running them in parallel easily causes mix-ups, and the saved results end up overlapping each other.


Strategy 2 - Fully parallel subagents
This is probably my favorite strategy, because each subagent is free to do its own work without depending on any earlier result, which pushes the processing speed to the maximum.
The requirement is that the task looks like processing independent files, where each result is saved in its own place, so there’s no risk of overwriting each other.
To make it easier to picture, say you have 12 Excel sales files from 12 branches. Each file names its columns differently, dates are sometimes dd/mm/yyyy and sometimes mm/dd/yyyy, and product names are a mix of upper and lower case. The main agent splits the work among 4 subagents, each cleans 3 files and saves them to a separate folder for each branch. Since no file depends on another, the total time is roughly the time it takes to process 3 files instead of all 12. Finally, the main agent merges the cleaned files and checks that the row counts match the original data.
The prompt the main agent gives each subagent could look like this, with only the file list changing between subagents:
You clean the sales data in the files assigned to you.
1. Read the cleaning rules in rules/clean_rules.md (column names, date format, product names).
2. Process these files one by one: data/raw/branch_01.xlsx, branch_02.xlsx, branch_03.xlsx.
3. Save each cleaned file to data/clean/<branch name>.csv, and never write to another subagent's files.
4. If you are not sure how to fix a row, keep it as is and log it in data/clean/<branch name>_review.csv.
When done, reply briefly: row counts before and after for each file, and how many rows need review.The trade-off is that running in parallel uses more tokens, because every subagent has to read the rules from the start. The prompt also has to be consistent, otherwise each subagent will process the data its own way.


Strategy 3 - Parallel subagents combined with sequential processing
This is a combination of strategy 1 and strategy 2, with both parallel and sequential parts. Going back to the group chat classification task, if there are many group chats, we can assign each subagent just one batch from a single group chat, and each group chat is still processed batch by batch in order.
Many tasks have this shape. The key point when splitting work among subagents is to make sure the data each subagent processes and saves is independent of the others, for example by splitting by customer ID.


Optimize your parallel processing skill
You can ask the agent to create a shared skill for all your projects, as long as the agent knows how to pick the right strategy for your task. I think having a global skill like that is worth it. But for projects with special, complex processing, you should also create a dedicated skill for that project, for these reasons:
- Split the work automatically with a script: Since the skill belongs to the project, its script can split the work based on the project’s data and the conventions between you and the agent. This makes every split consistent. For example, the script checks how many group chats still have unprocessed messages and prints a report, so the main agent knows which group to assign to each subagent.
- Consistent prompts for the project: The project skill holds the prompt or brief made specifically for this work. This makes sure every time the main agent calls a subagent, it uses the same format.
The skill should include these parts:
- Rules for splitting work among subagents and which strategy to use.
- Rules for how subagents return results: how to report, and which folder to save results in so each subagent’s output is easy to tell apart.
- How the main agent combines the results and reports back.
In short, if your data depends on each other, run subagents sequentially; if it’s independent, run them in parallel; and if you have both, combine the two. If you have a better way to use subagents, feel free to share it with me in the comments.






