Establishing a rolling research program for Google Beam

Timeline

Oct 2025 – Apr 2026

Role

UX Researcher

Org

Google

Methods

Rolling research, Moderated studies, Usability Metrics

MY MISSION

I helped the Beam team learn continuously before launch

I ran 5 research cycles over 6 months, helping Beam teams evaluate evolving product experiences, compare signals across studies, and make pre-launch design and engineering decisions.

Beam introduced me to an environment where research was already deeply embedded into product development. My role became learning how to operate within that system, scale its impact, and help teams move confidently from questions to decisions.

LET ME TELL YOU ABOUT BEAM

Beam is Google's attempt to make remote conversations feel less like a video call and more like sitting across from another person.

Instead of focusing only on audio and video quality, Beam focuses on something harder to measure: the feeling of being present with someone. That meant our research wasn't just evaluating whether features worked, it was evaluating how interactions felt.

WHAT I DID

This is what the rolling research program looked like..

The work was not just running individual studies. it was building a repeatable system that helped teams decide what to learn, how to evaluate it, and how to carry evidence into the next cycle.

..where research included exploring and validating..

EXPLORE NEW FEATURES AND INTERACTIONS

Understanding how new interactions performed against Beam’s key value indicators.

VALIDATE EXISTING FEATURES AND INTERACTIONS

Stress-testing how existing software and hardware interactions perform before launch.

…using qual for the 'why' and quant for the 'how many'?

The qualitative sessions helped explain why an interaction felt natural, awkward, or confusing. Standardized measures helped us compare signals across cycles so the team could see whether the experience was improving, regressing, or changing in unexpected ways.

STUDY PREP

Before the Study

Before each lab cycle, i worked through the decisions that would make the sessions useful: what we needed to learn, how we would capture it, and how the data would stay comparable across studies.

WHAT SHOULD WE LEARN?

I used the calendar to turn open questions into study priorities

I maintained a public rolling research calendar and an open question repository, transforming our intake from a list of requests into a collaborative prioritization session.

01

Created visibility into upcoming monthly research cycles across teams.

02

Captured evolving product questions as they emerged instead of waiting for kickoff meetings.

03

Turned research intake into a continuous collaborative process.

04

Made prioritization proactive instead of reactive.

For example, the design team wanted to know,

Open questions:
"We're introducing Presentation Mode. Does it help people stay focused on the presenter without making the meeting feel less natural?"

It translated to the study as:

Observe how participants interact with, how they interpret it, and whether it changes conversational flow during a collaborative task.

Or, the engineers asked,

Open questions:
"If joining takes a few extra seconds, is that something people even notice, or does it only become a problem after a certain point?"

It translated to the study as:

Measure participants' expectations during the join experience, identify thresholds where waiting becomes noticeable, and capture behavioral and verbal reactions as delays increase.

HOW SHOULD WE CAPTURE IT?

I translated the study plan into Qualtrics for data collection

I set up Qualtrics as the central workspace for translating the study plan into structured tasks, post-task reflections, standardized ratings, and moderator-facing capture fields.

01

Created a consistent place to capture post-task reflections, ratings, and open-ended feedback.

02

Standardized how recurring experience qualities were measured across study cycles.

03

Made it easier to compare signals across participants, scenarios, and monthly rounds.

04

Reduced analysis friction by keeping qualitative notes and quantitative measures organized from the start.

For example, "Does Presentation Mode help people stay focused on the conversation?"

Behavior:
Did participants continue interacting naturally during the task?

Rating

Presence (1–7) and Ease of collaboration (1–7)

Reflection

"Did anything interrupt the flow of the conversation?"

Moderator notes
Moments of hesitation, attention shifts, or conversational repair.

METRICS

I built component-level metrics for Beam’s design system

Beam already had baseline KVIs for the overall experience. my work focused on the next layer: helping design systems understand how specific UI components and interaction patterns contributed to those broader experience goals.

01

Mapped components to experience goals: Connected UI patterns back to Beam’s existing KVIs so component decisions could be evaluated against product-level experience outcomes.

02

Defined component-level success criteria: Clarified what “good” meant for specific components beyond visual consistency or technical correctness.

03

Standardized reusable measurement prompts: Created rating and reflection language that could be reused when similar components appeared across studies.

04

Made component impact comparable across cycles: Helped design systems partners understand whether component changes were supporting or weakening the broader Beam experience over time.

STUDY SETUP

The sessions had to feel like real conversations, not isolated tasks

Beam is about presence, the session design had to preserve the social context around each interaction. the goal was to see how people behaved with another person in the room, not just whether they could complete a task.

AI collaboration in planning

I used AI to pressure-test discussion guides and identify places where prompts might bias participants toward describing “presence” too abstractly. Final protocols were reviewed against stakeholder questions and the product context before sessions.

Click to jump

DYADIC SETUP

Pairs helped us see what solo testing would miss

I guided participants through realistic meeting prompts while avoiding over-directing the interaction. The goal was to keep the session structured enough to compare across studies, but open enough for natural conversational behavior to emerge.

Structured prompts kept scenarios consistent across participants and cycles and light-touch moderation allowed natural hesitation, repair, and social behavior to surface.

NDA-FRIENDLY EXAMPLES

"You're reviewing a presentation with your partner. Work together to identify five differences between the slides."

"Review this document together and decide which edits should be accepted."

SESSION FLOW

The flow stayed consistent, even when the questions changed

Each session combined baseline metrics, realistic scenarios, post-task ratings, open feedback, and a final debrief. This gave us a consistent way to compare signals across studies while still capturing the qualitative nuance of live conversation, including attention shifts, hesitation, confidence, and moments where the interaction either supported or disrupted presence.

PARTICIPANTS

We needed both fresh eyes and people who knew Beam

I worked with the recruiting team to source paired participants who matched Beam’s target enterprise audience, guided by criteria such as role, company context, collaboration habits, meeting tool usage, and prior familiarity with Beam.

Including both new and returning Beam users helped us understand where the experience needed to be immediately legible, and where feedback reflected deeper expectations from people who had already used the product. This distinction helped us separate onboarding friction from deeper signals about interaction quality, presence, and product expectations.

STUDY ANALYSIS

After the Study

After each lab cycle, I brought together post-task ratings, moderator notes, participant reflections, and observed behavior to understand what was recurring, what was still ambiguous, and what teams could act on before launch

WHAT WERE WE SEEING?

I shared early signals before waiting for a final report

After each session, I reviewed post-task ratings alongside qualitative notes, participant quotes, and observed behavior. I then shared lightweight research memos so design, product, and engineering could understand what we were seeing before the full cycle ended.

01

Compared post-task ratings with observed behavior to understand what was driving each score.

02

Separated repeated signals from isolated moments so the team could avoid overreacting to one session.

03

Flagged areas where confidence was emerging, mixed, or still too early to call.

04

Shared early memos that helped teams act on directional learning before the next research cycle.

NDA-FRIENDLY EXAMPLES

Early signals we're seeing on connection quality

  • Mild degradation often went unnoticed.

  • More severe degradation interrupted turn-taking rather than task completion.

  • Participants adapted quickly, but conversations became noticeably less fluid.

  • Need another round to understand where the tipping point occurs.

WHAT SHOULD WE DO NEXT?

I organized findings by what teams could do with them

I organized insights into a lightweight framework that made the evidence easier to act on. Instead of treating every issue equally, I helped teams understand which signals were most important, how confident we were, and what kind of product decision each insight could support.

01

Mapped each insight to the product question or decision it could inform.

02

Tagged signals by confidence level, based on recurrence, severity, and supporting evidence.

03

Separated usability friction from deeper presence-related signals like attention, timing, and social comfort.

04

Turned findings into next-step recommendations for design iteration, follow-up research, or launch readiness.

IMPACT

The research informed launch decisions across 5 studies

Across six months, findings from the rolling studies helped design and engineering decide what to refine, validate, or carry forward before launch.

Design team adopted the metrics work as a shared measurement language

The repository helped teams evaluate component changes more consistently across studies and connect UI decisions back to Beam’s broader experience goals.

Some kind words from my research lead!

"I brought Ambika onto my team to support research for Google Beam. Ambika was instrumental in establishing and scaling our rolling research model, a structured program that continuously generated insights to support our fast-moving product decisions. Her ability to run these research cycles at pace was critical in ensuring our design and engineering teams remained data-informed throughout the development process. Ambika did a great job of bridging the gap between research and product direction. She effectively connected insights from her research to team priorities and communicated them clearly to a wide range of stakeholders across design, engineering and product management. I can recommend Ambika as a great researcher who can perform disciplined research execution while clearly communicating insights."

Aaron Rich
Aaron RichUX Research Lead, Google Beam

The detailed insights are confidential.

Please email me if you'd like to chat!

Where to Find Me