Shopify Storefront Agent Evaluator
Mercor • Remote
Education
Any
Type
hourly
Pay Rate (by country)
$40–$50/hr
Listed
Today
✅ Applying through this link supports our platform at no cost to you.
This position is hosted on an external talent platform. Please only apply for this position if it fits your skills and interests.
In our Talent Pool?
Apply through this link and we can vouch for you to Mercor. ? We vouch for Talent Pool members who apply through our referral link, when we believe they're a strong match. Not every applicant gets a vouch. Not in the pool yet? Set up your profile first.
Set up your profile →Mercor: our referral track record
We've referred 556 candidates to Mercor roles. 5% (29) were placed.
What We Know About This Role
- Weekly hours
- 5 hrs/week
About this Role
From the Mercor listing
About the role
Shopify's Storefront Agent is the AI shopping assistant that sits on a merchant's store — it searches the catalog in natural language, recommends products, builds carts, and answers questions about shipping, returns, and store policies.
Your job is to stress-test it and grade what comes back. You'll send the agent a handful of realistic shopper prompts, then score each response against a short rubric. The work is straightforward and self-contained: no coding, no data pipelines, no long-form writing.
What you'll do
Prompt the Storefront Agent as a real shopper would — product discovery, comparisons, sizing and stock questions, shipping and returns, edge cases and awkward requests
Grade each response on a small set of dimensions:
Relevance — did it actually answer what was asked, within the constraints the shopper gave?
Accuracy — is every claim (price, stock, policy, delivery window) true and grounded in the store's own catalog and policy data, rather than plausible-sounding invention?
Safety — did it avoid harmful advice, unsupported claims, and actions it wasn't authorized to take?
Leave a short, concrete note on anything that failed, so the reason is legible to someone who wasn't there
Commitment: up to 5 hours/week
Qualifications
Required:
Consumer e-commerce fluency — you shop online regularly and have a clear sense of what a good vs. a useless product recommendation looks like.
Careful, consistent judgment — you can apply the same rubric the same way across many responses, and separate "this response was wrong" from "this response wasn't what I'd have written."
Clear, concise written English — enough to explain in a sentence or two why a response failed.
Strongly preferred:
Hands-on Shopify merchant experience — you've run or operated a Shopify store and know how shoppers actually behave on one.
Experience using AI shopping assistants or chat agents as a customer, and a feel for where they tend to break.
Preferred (nice to have — we'll ramp you on the specifics):
Prior work evaluating, red-teaming, or annotating AI model outputs.
Familiarity with retail operations: catalog and variant structure, inventory, shipping and returns policy.
What this is not
This is not a customer-support role and not an engineering role. You are not fixing the agent — you are judging it, precisely and repeatably, so the team can measure where it falls short.
Requirements
- Must be eligible to work in Remote
- Fluent proficiency in English (Written & Verbal)
- Reliable high-speed internet connection
- Bachelor's degree or equivalent professional experience
- Demonstrated expertise in Generalist
Eligible Languages
Fluent proficiency in English
Key Responsibilities
- Prompt the Storefront Agent as a real shopper would — product discovery, comparisons, sizing and stock questions, shipping and returns, edge cases and awkward requests
- Grade each response on a small set of dimensions:
- Relevance — did it actually answer what was asked, within the constraints the shopper gave?
- Accuracy — is every claim (price, stock, policy, delivery window) true and grounded in the store's own catalog and policy data, rather than plausible-sounding invention?
Why This Role
Work from anywhere, at any time. This fully remote Shopify Storefront Agent Evaluator position ($45/hr) breaks down geographic barriers, allowing you to earn US-competitive rates regardless of your local market. It is a perfect stepping stone for building a career in the Generalist ecosystem.
Skills & Categories
Explore other opportunities in related specializations:
Related Jobs
Browse All Jobs from Mercor
Discover more opportunities on Mercor that match your skills and interests.
View All Mercor Jobs →Verified Reviews
Community Reviews
Share your experience with Mercor
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Frequently Asked Questions
Is Mercor for freelancers or full-time contractors?
Mercor places you with one client for a defined engagement, like 'Python Tutor for 3 months', rather than having you grab small tasks from a shared queue. Most roles function as steady contract work, not one-off gigs.
Does Mercor's application require an on-camera interview?
Yes, every applicant records a video interview with an AI interviewer that asks questions about your resume. Clients review that recording to judge communication skills before matching, so there's no way to apply without going on camera.
Does it cost money to apply to Mercor?
No, applying and joining Mercor is free. Mercor's revenue comes from a fee it charges the client on top of your hourly rate, not from applicants. Treat any request for payment to join as a red flag.
What does task-based AI training work actually look like?
Practical, hands-on data work: recording short videos, categorizing images, rating text responses, or analyzing data. Tasks are designed to be short and distinct, typically 5 to 60 minutes each.
What does asynchronous AI training work mean in practice?
No set hours, no check-ins, no meetings. You log in when you want, pick up an available task, complete it, and submit; nobody is waiting on you in real time. That's different from remote employment, where you're expected online during business hours. The tradeoff: you're competing with others for available tasks, so an empty queue means there's simply nothing to do until more work is released.
What does Generalist work look like for a Shopify Storefront Agent Evaluator?
Tasks here are scoped to Generalist, not generic labeling. As a Shopify Storefront Agent Evaluator, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Generalist) rather than following a one-size-fits-all rubric. If you don't have hands-on Generalist background, this is likely not the right listing to start with.
How many hours per week does this role require?
Based on the listing, this role is scoped at about 5 hours per week. Treat this as a real commitment expectation, not a loose estimate.
Do I need to be fluent in English?
Yes. This role specifically requires English proficiency. You will likely be evaluated on written fluency during the assessment, not just conversational level. If English is not your first language or you are not professionally fluent, this is not the right role. Filter for your native language to find better-matched listings.
What happens when I click Apply on this listing?
You'll be taken to Mercor's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.
How soon will I start working after applying to Mercor?
Not immediately. Mercor is a talent marketplace, not a task queue, so applying puts you in a pool of candidates. You start working only once a specific client, like a major AI lab, selects your profile, and that matching process can take weeks.