Creative testing
Structured creative testing for beauty, ecommerce, and consumer brands. Disciplined experiments across hooks, formats, and messaging that find what actually moves the needle.
Testing creative without structure is just spending. We run a disciplined creative testing program across Meta, TikTok, Google, and YouTube that isolates variables, hits statistical significance, and files every learning where the whole team can use it.
The objective isn't activity. It's walking out of every quarter with a proven creative playbook the brand didn't have three months earlier. That's how we've helped beauty and consumer brands like Sol de Janeiro compound performance quarter over quarter instead of leaning on guesswork.
We test creative with a point of view, not just variations
Great creative isn't random, so our tests aren't either. We test against the two story structures that actually earn attention, and we know which one a placement calls for before we build a single variant.
The Hill is the traditional arc: lead in, build, climax, offer, branding. It works when you have someone's attention for thirty seconds and a clear payoff to deliver.
The Heartbeat
The Heartbeat is built for the feed: start high, seed subtle brand cues, break the pattern with an unexpected shift, stack multiple peaks, and save more story for the people who keep watching. It's built for a world where the first two seconds decide everything.
Every test starts from a hypothesis about which arc, which hook, and which format the audience will reward, so a win teaches us something we can use again instead of a lucky ad we can't explain.
What we do
- Test design and prioritization. Every creative test starts with a hypothesis, a KPI, and an expected P&L impact. We keep a prioritized backlog so spend goes to the hooks, formats, and offers worth answering, not whatever is easiest to launch.
- Statistically sound readouts. Sample sizes, holdouts, and confidence thresholds are agreed upfront, so results get trusted instead of debated. We call tests when the data says to, not when the calendar does.
- Learning libraries that scale. Winners feed the creative playbook and losers become guardrails. Every result is documented by brand, platform, and format, so the team never re-runs the same mistake, and new launches start from what we already know instead of zero.
How we find a winner
We run creative like a bracket. Champions, the ads currently carrying performance, go up against challengers, and only the winner advances to the next round. We start with the highest-impact elements, hook and format and offer, then drill into the smaller variables once the big questions are settled. Nothing earns budget on opinion. It earns it by beating what's already working.
How we run it without touching core performance
Tests are flighted, not thrown live all at once. Every flight has a clear window and a defined budget, mapped out so the whole team has visibility and, more importantly, so a test never quietly drags down the account's core performance. While a test is running, the hypothesis, spend, and performance notes live in one tracker instead of a Slack thread, so a mid-flight read is a glance rather than a fire drill. We map the testing agenda across the quarter, expansion tracks, evergreen refreshes, net-new variables, and seasonal pushes like holiday, so learnings compound in the right order instead of colliding.
Our approach
Every test has a hypothesis, a KPI, and an expected impact. We test what could actually move the P&L, not what's easy to launch.
Sample sizes, holdouts, and confidence thresholds are agreed upfront, so results get trusted, not debated. No peeking, no moving the goalposts.
Winners feed the creative playbook. Losers become guardrails, so the team stops re-running the same mistake.
Outcomes
- Compounding CAC efficiency quarter over quarter, because every test result, win or lose, permanently improves the playbook instead of evaporating when the campaign ends.
- Faster launches into new markets and new products, because hooks and formats already proven for the brand carry over instead of getting rediscovered from scratch.
- A shared creative playbook that informs landing pages, email, and the next round of creative briefs.
FAQ
- How long does a creative test usually run?
- At healthy spend, most tests read in about a week to ten days. We agree on the read window before we hit publish, so no one gets tempted to call a winner three days in.
- How much budget do we need for a real testing program?
- A good rule of thumb is putting 10% to 20% of your paid media budget toward testing. Under about $50K a month, we run fewer but bigger tests. Above that, we can run parallel creative tracks without stepping on each other.
- What do you actually test?
- Hooks, formats (static, video, UGC, motion), messaging angles, and offers. What jumps to the top of the queue is whatever we think can move the P&L most. Novelty on its own doesn't earn a slot.
- Do you produce the creative for the tests?
- Yes. Our creative team builds the variants (statics, video, UGC-style) straight from the testing backlog. When brands have their own in-house teams or outside partners, we're glad to work alongside them too.
- How do you keep tests from hurting overall performance?
- We flight them. Each test gets a defined window and budget, mapped so the full team has visibility and no single experiment can drag down the core account. Champions keep carrying performance while challengers earn their way in.
- What happens to the learnings after a test ends?
- Every result goes into a library sorted by brand, platform, and format. Winners shape the next briefs and media plans. Losers become guardrails, so we stop chasing angles that already lost.

