Concepts / User Feedback in Web Services

User Feedback in Web Services

Personalization uses a user's online activity to infer a profile of interests and preferences.

  • Programming

From One Choice to a Personal Choice

A web service often has to decide what to show a user, such as which news article or advertisement to deliver. Without personalization, that decision can be a single choice for everyone. Personalization uses the user's online activity to infer a profile of interests and preferences, allowing the service to make a choice guided by information about that particular user.

The central idea is a continuing loop: the service selects content, the user responds through activity or feedback, and that new information can influence later recommendations.

informsguidesselectselicitsinfluencesadjustsOnline activityUser profileInterests and preferencesRecommendation policySelects contentSelected contentArticle or advertisementUser feedbackSurvey or clickPolicy improvementLater recommendations
How does user activity flow through profile inference, content recommendation, feedback collection, and policy improvement?

How a Profile Guides Content Selection

A user profile is the link between observed activity and the policy's content choices. The service uses online activity to infer interests and preferences. A recommendation policy then uses that profile to select content for the particular user. The selected content might be a news article or an advertisement.

guidesmay selectmay selectUser profileInterests and preferencesRecommendation policyChoice guided by profileNews articleAdvertisement
How does a recommendation policy use a user's profile to choose which content to present?

Choosing Content for a Particular User

Trace the role of a profile when a web service decides what to deliver.

Observe activity: The service uses the user's online activity as information about the user's interests and preferences.

Infer a profile: That activity is used to form a user profile containing inferred interests and preferences.

Apply the policy: The recommendation policy uses the profile to guide its content choice for this user.

Deliver content: The service delivers selected content, such as a news article or an advertisement.

The profile connects user activity to the recommendation policy's choice of content.

Two Sources of User Feedback

Feedback sourceWhat the service observesHow it can be used
Satisfaction surveyA response deliberately provided by the userProvides direct information about the user's satisfaction
User clickA click on a link monitored by the serviceActs as an indicator of interest in that link

These two feedback sources differ in how the information is obtained. A satisfaction survey asks the user to provide feedback directly. A monitored click is behavioral feedback: the service treats the user's action as an indicator of interest in a link. Both can provide information that influences later recommendations.

From Feedback to Policy Adjustment

Reinforcement learning can improve a recommendation policy by adjusting it in response to user feedback. In this setting, the policy selects content, the user responds through activity or feedback, and that response becomes information about how the choice should influence future recommendations. The policy is therefore not treated as a one-time, unchanging rule.

choosesproducesinformsinfluencesInitial policySelects contentContent choiceArticle or advertisementUser responseSurvey or clickReward signalFeedback used for learningAdjusted policyFuture content choices
How does a user's feedback become a reward signal that changes the recommendation policy over time?

Tracing One Feedback Cycle

Follow the information flow when a service presents content and then observes a user response.

Select: The recommendation policy uses the user's profile to select content for that user.

Observe: The service collects feedback through a satisfaction survey or by monitoring whether the user clicks a link.

Interpret: The survey response or click provides feedback about the user's response to the selected content.

Adjust: Reinforcement learning can use the feedback to adjust the recommendation policy.

Recommend later: The adjusted policy can influence later content recommendations.

Feedback closes the loop between a current content choice and later policy behavior.

Mistakes in Tracing the Loop

  • Treating personalization as a one-time transfer of information

    Personalization is described as a two-way process. The service sends selected content, and the user's activity and feedback can influence later recommendations.

    Fix: Trace both directions: profile-guided content selection and feedback flowing back toward later recommendations.

  • Confusing the user profile with the recommendation policy

    The profile contains inferred interests and preferences. The recommendation policy uses that profile to select content.

    Fix: Describe the profile as the link that guides the policy's choice.

  • Listing only surveys as feedback

    The source identifies both satisfaction surveys and monitored clicks as ways to collect feedback.

    Fix: Include direct survey responses and clicks that act as indicators of interest.

  • Inventing a specific policy-update formula

    The source says that reinforcement learning adjusts the policy in response to feedback, but it does not specify a mathematical scoring method or adjustment rule.

    Fix: Explain the information flow without assigning an unsupported formula or fixed update amount.

Apply the Information Flow

MEDIUM

A web service uses a user's online activity to infer interests and preferences. It then selects an advertisement and monitors whether the user clicks it. Trace the process in four stages: identify the profile information, name the component that selects the advertisement, identify the feedback, and explain what reinforcement learning may do with that feedback.

Hints
  • Separate the inferred profile from the recommendation policy.
  • A monitored click is one of the feedback sources identified in the lesson.
  • State that the feedback can influence an adjustment without inventing a numerical update rule.

What do you think happens?

Which description best traces the feedback loop?

  • The service selects content once and ignores the user's response.
  • The service infers a profile, uses a policy to select content, collects feedback, and can adjust the policy for later recommendations.
  • The user profile directly delivers content without a recommendation policy.
  • Only written surveys can influence recommendations.
Reveal answer

Answer: The service infers a profile, uses a policy to select content, collects feedback, and can adjust the policy for later recommendations.

The profile links observed activity to the policy's content choice, while surveys and monitored clicks provide feedback that can influence later policy adjustments.

Key Takeaways

  1. Personalization uses online activity to infer a profile of interests and preferences.
  2. A recommendation policy uses that profile to select content for a particular user, such as a news article or advertisement.
  3. A web service can collect feedback through satisfaction surveys or by monitoring clicks as indicators of interest.
  4. Reinforcement learning can adjust the recommendation policy in response to feedback.
  5. The profile and policy form part of a continuing loop in which user activity and feedback can influence later recommendations.

Key Takeaways

  • Personalization turns online activity into a profile of inferred interests and preferences.
  • The recommendation policy uses the profile to choose content for a particular user.
  • Satisfaction surveys and monitored clicks are two ways to collect user feedback.
  • Reinforcement learning can use that feedback to adjust the policy and influence later recommendations.