Skip to content

E-commerce

AI in e-commerce – personalisation, recommendations and automation

Read the articleQuestions and answers

Article cover: AI in e-commerce – personalisation, recommendations and automation
AI in e-commerce delivers real results only when it is based on complete, consistent data and clearly defined rules for its use. Personalisation and recommendations can increase offer relevance in real time, but they require proper tracking of behaviour and well-structured product descriptions. In practice, most mistakes do not come from a “weak model”, but from issues on the event side, user identification and catalogue quality. Just as important is the implementation layer: a data warehouse for training and, where needed, streaming and cache for rapidly refreshing signals. At the same time, personalisation must take privacy, marketing consents and access control into account; otherwise, you expose yourself to lower trust and compliance-related risks. Further on, you will find the minimum data set and key privacy/compliance requirements without which AI does not deliver results.

What data is essential for effective personalisation in e-commerce?

For effective personalisation in e-commerce, you need at least behavioural data, product catalogue data, stock data, price history and customer data in the form of identifiers and segments. The foundation is events such as view/add_to_cart/purchase and a consistent product_id, because errors such as double sending or missing purchase can throw even carefully designed recommendations off track. On top of that, you need a catalogue with price, margin and attributes. In fashion, size/colour/brand are critical, and in electronics, technical parameters and compatibility. If you are planning pricing, without data on costs and gross margin per SKU as well as return costs, it is easy to increase turnover at the expense of profit.

  • Behavioural events: view, add_to_cart, purchase
  • Product catalogue: price, margin, attributes (e.g. size/colour/brand or technical parameters)
  • Stock levels and price history
  • Customer data: anonymous ID, segments, acquisition channel

Scalable personalisation also requires consistent user identification across devices and a well thought-out data processing architecture. Most often, an ID in the analytics layer is used (e.g. GA4 + BigQuery, Amplitude, Segment) together with linking user_id after login with anon_id from cookies, along with merge rules after authorisation. It is worth maintaining 1–2 canonical identifiers, because “ID chaos” lowers model quality and underreports attribution. On the infrastructure side, a data warehouse (BigQuery/Snowflake/Databricks) is usually enough for training, while Kafka/Pub/Sub and cache (Redis) with a delay of around <1–5 minutes come in handy for updating signals in near real time.

For model results to remain stable in production, data quality and feature consistency mechanisms are needed. In practice, this means implementing validations in ETL (dbt tests), bot filters (e.g. Cloudflare Bot Management) and compliance controls, where the number of purchases in events should match ERP/OMS within a margin of, for example, 1–2%. For online/offline consistency, a feature store (e.g. Feast, Tecton) is used so that identical feature definitions apply in training and inference, and training-serving skew does not occur. When you care about understanding product similarity across categories and naming, embeddings are used and stored in vector databases (Pinecone, Weaviate, Milvus, pgvector), which supports search and recommendations.

Personalisation What data is essential for effective personalisation in e-commerce?
  1. 01User behaviourViews, add to cart, purchases (events).
  2. 02Catalogue and stockConsistent ID, price history, availability.
  3. 03Identification and segmentsConsistent identification, cross-device, segments.
  4. 04Pricing and marginCosts, gross margin, return costs.

The foundation of effective scalable personalisation is consistent data on behaviour, catalogue, identification and transaction costs, without which recommendations can lead to errors and losses.

What are the privacy and compliance challenges in AI personalisation?

The biggest challenge in the area of privacy and compliance in AI personalisation is the secure processing of customer data while at the same time respecting consents and proper access control. In practice, pseudonymisation is used (e.g. hashing email addresses) and PII is separated from behavioural events to reduce the risk of unauthorised access. At the same time, permission management is implemented (IAM, row-level security in BigQuery/Snowflake) so that data/AI teams work only on data they are entitled to use. A good practice is auditing: who ran the pipeline, on which version of the data and model, and whether marketing consents were respected.

The second critical area remains consent management and the “right to be deleted”, because personalisation cannot operate “alongside” compliance policies. Integration with a CMP (e.g. OneTrust) is needed, along with a mechanism that enforces consent compliance in data pipelines. “Right to be forgotten” is also implemented, meaning deletion of user data from the data warehouse and feature store within a defined SLA (e.g. 30 days), as well as anonymisation of identifiers in logs. This allows personalisation and marketing automation to be based on reliable signals, without the risk that the system uses data for which the customer did not give consent.

How can AI improve the shopping experience through personalisation?

AI can improve the shopping experience by adapting content and product placement order to the user’s context and behaviour in real time. In practice, this means onsite personalisation, where on the PLP you can sort products by preferred brands or colours, and adapt the hero banner to the traffic source (e.g. Google Shopping vs direct), while keeping a limit on changes so the UX remains stable. Context (time of day, device, channel, location) increases relevance when behaviour differs between channels, e.g. on mobile price and shorter sessions matter more. The biggest advantage of on-site personalisation is that it covers all traffic — including anonymous users — and affects purchase decisions “here and now”.

AI also improves the off-site experience by personalising communication in e-mail/SMS/push in terms of content, send time and frequency. In tools such as Klaviyo, Braze or Salesforce Marketing Cloud, it is possible to combine product recommendations with rules that limit communication fatigue (frequency capping), e.g. a maximum of 2 SMS a week. In addition, send-time optimisation makes it possible to adjust timing, while personalising the order of products and arguments increases the chance that the message will be useful rather than sound “mass”. This combination of controls (rules) and tailoring (AI) makes it easier to maintain a consistent customer experience across multiple channels.

The most “noticeable” personalisation appears when you automate entire scenarios rather than just individual messages. Customer journey orchestration connects events (e.g. abandoned basket, drop in activity, purchase) with a decision on the “next best action” and channel selection (e-mail vs push vs remarketing). For example, if a customer abandons their basket and has a high propensity to buy, you send a reminder without a discount, and if it is low — you test a 5% voucher with a 24h limit. As a result, the user receives communication and an offer tailored to the situation, rather than a random campaign sequence.

AI personalisation How does AI improve the shopping experience?
  1. 01Real-time contextTime of day, device, location
  2. 02On-site personalisationTailoring content and sorting
  3. 03Stable UX experienceA limit on changes for comfort
  4. 04Off-site communicationE-mail and SMS

AI tailors content and the order in which products are displayed to the real-time context, influencing purchase decisions “here and now”.

Recommendation systems – from “similar products” to personalisation

Recommendation systems evolve from simple “similar products” modules to 1:1 personalisation when you precisely define the goal and measure the impact on sales and profitability. CTR can be misleading because it favours “clickable” propositions, so KPI such as CVR, RPV (revenue per visitor) and the impact on AOV and margin are more often chosen. In practice, for high-margin categories, the ranking is sometimes calculated so that it takes into account both the probability of purchase and the expected margin (e.g. 70% purchase + 30% margin). If recommendations are to deliver business results, their KPI should reward not only clicks, but also revenue per user and margin.

The usual starting point is an approach that delivers results immediately, and only then do you add behavioural signals once interaction history has been collected. Content-based (similarity of attributes and text) works well from day one, whereas collaborative filtering (e.g. implicit ALS) requires data on clicks and purchases. For logged-out users, where only the session is visible, sequential models are used to predict the “next product” based on the last 5–20 actions (e.g. GRU4Rec, SASRec, Transformers for sequences). On top of that, there are contextual recommendations that take into account differences between channels, e.g. mobile vs e-mail.

The return from recommendations depends heavily on their placement on the site and on how the product list is arranged. On the PDP, the “similar” and “frequently bought together” sections usually work well; in the basket, cross-sell with high margin and low return risk makes sense; and on the home page, category and brand personalisation performs better. To avoid damaging UX, it is usually worth limiting the number of modules (e.g. 2–3 per page) and keeping API render latency in check (ideally <100–200 ms). At larger scale, ranking and re-ranking come into play: first selecting candidates, then arranging the top results with availability and margin goals in mind (e.g. XGBoost/LambdaMART re-ranker).

Full personalisation also requires handling cold start and robust testing, because “nice” offline results do not have to translate into growth in production. For a new user, contextual popularity (top in category, 24h trends) and a short preference questionnaire are used, and for a new product — an embedding from the description/attributes and similarity to existing SKUs. Quality is assessed offline (MAP@K, NDCG@K, Recall@K) with time-based validation, but deployment is determined by online tests (A/B or bandit) based on business KPIs. The minimum standard for an online test is 2–4 weeks or reaching statistical power, e.g. detecting +2% RPV at a significance level of 0.05.

Marketing automation and communication personalisation with AI

Marketing automation with AI comes down to matching segments, content and the time of contact to customer behaviour so that campaigns are profitable, not just “nice”. A common starting point is RFM segmentation (Recency/Frequency/Monetary), while a more advanced approach uses CLV and churn prediction to direct budget towards people with real potential. In practice, this makes it possible to design win-back scenarios with a discount cap rather than launching blanket promotions for everyone. Such segmentation also makes it easier to control promotion costs and maintain a consistent communication strategy.

Personalisation in e-mail/SMS/push delivers the best results when AI selects not only the products, but also the timing and frequency of contact. Tools such as Klaviyo, Braze or Salesforce Marketing Cloud make it possible to combine product recommendations with rules that limit communication fatigue (frequency capping), e.g. a maximum of 2 SMS a week. Send-time optimisation and frequency control are practical “levers” that increase the usefulness of communication without raising promotional pressure. This means automation supports sales while also reducing the risk of overexposure to messages.

Generating creatives and copy through language models makes sense provided there is strict control over the brand and the facts, which is why in practice templates and data-driven generation based on PIM data (e.g. Akeneo) are used, rather than relying on a “pure prompt”. An LLM can prepare variants of product descriptions or ad headlines, but the final text undergoes validation for compliance with parameters and for the prohibition on making false promises or attributing non-existent certifications. In paid campaigns, the biggest AI lever is often feed quality (titles, attributes, GTIN, images) and first-party signals, which improve matching and attribution. To reliably assess the impact of personalisation, tests with random assignment to variants and one primary metric (e.g. RPV) are used, and the incremental effect is also verified through holdouts.

Automation & personalisation Marketing automation and communication personalisation with AI
  1. 01Advanced segmentation (RFM & CLV)Prediction of potential and churn risk.
  2. 02Effective win-back scenariosPersonalised offers with a discount cap.
  3. 03AI personalisationOptimal timing, frequency and content.

Key: focusing the budget on customers with real potential for profitable campaigns.

Operational automation in e-commerce (customer service, logistics, catalogue) with AI

Operational automation with AI streamlines customer service, logistics and catalogue work when models make repetitive decisions based on data and business rules. In customer service, the safe pattern remains RAG (Retrieval-Augmented Generation), in which a chatbot/voicebot answers based on documents (terms and conditions, FAQ, order statuses) and can point to sources. Deployments of this kind use, among others, Zendesk AI, Intercom Fin, Freshdesk or a custom solution on Azure OpenAI/Vertex AI with searching in Elastic/OpenSearch. By grounding responses in a specific knowledge base, the risk of incorrect communications is significantly reduced.

Ticket automation shortens handling time when AI classifies enquiries, assigns priorities to them (e.g. VIP, delayed parcel) and suggests response drafts to agents. In practice, it is better to measure the effects using operational metrics such as AHT (average handle time), FRT (first response time) and CSAT after contact, rather than judging the system only “by feel”. The greatest value comes from linking routing and priorities with response suggestions, because it reduces the number of escalations without increasing headcount. This approach organises the team’s work and improves service predictability during periods of higher traffic.

In logistics and operational finance, AI supports decisions on stock, delivery and risk, provided the right input signals are available. Inventory forecasting combines sales history, seasonality, campaigns and lead time to calculate the reorder point and safety stock per SKU, and with long lead times (e.g. 30–60 days) more conservative thresholds and “best/expected/worst” scenarios are used. When selecting a carrier, the model can choose the “cheapest option that meets the SLA” based on data on on-time performance, costs, damage and returns per carrier and region. In fraud detection, the analysis covers, among other things, address consistency, account history, velocity, device fingerprints and basket anomalies, while tools such as Sift, Riskified or Stripe Radar support risk thresholds and the manual verification workflow.

  • PIM and catalogue: extracting attributes from descriptions and images, then standardising dictionaries for search and filters (e.g. “USB-C/Type C/USB Type-C”)
  • Search: integration of Algolia/Elasticsearch with semantic reranking based on embeddings and business rules (availability, margin)
  • Returns and complaints: forecasting returns per product/customer and triage (defect vs damage in transit), with automatic decision-making in simple cases up to a specified amount (e.g. <100 zł)
  • Back office: RPA (UiPath, Power Automate) + AI for reading invoices, checking price compliance and updating statuses in ERP/OMS

Customer segmentation and communication personalisation

Customer segmentation and communication personalisation with AI means that you select audience groups and the content and intensity of contact based on behavioural data, instead of relying on a single “mass” rule. The starting point is RFM segmentation (Recency/Frequency/Monetary), and the next stage is CLV and churn prediction, which means budget and discounts go to people with real potential. For example, the segment “high CLV + churn risk” may receive a win-back sequence with a discount cap, rather than a broad promotion for the entire database. This logic reduces wasted budget on low-value customers and organises the CRM strategy around profitability.

Personalisation in e-mail/SMS/push channels works most reliably when AI controls not only the recommended products, but also the timing and “frequency of contact” (frequency capping). In tools such as Klaviyo, Braze or Salesforce Marketing Cloud, you can combine product recommendations with rules, e.g. a limit on the number of messages and a “no discount” condition for customers with low price sensitivity. AI can also support send-time optimisation, i.e. choosing the send moment according to the recipient’s habits. In practice, the best results come from combining personalisation with guardrails, because the communication is tailored while still remaining under business control.

Logistics optimisation and delivery cost management with AI

AI optimises logistics and delivery costs when it automatically selects the carrier based on data on on-time performance, costs, damage and returns broken down by region. Instead of rigid rules, the system can choose the “cheapest option that meets the SLA”, e.g. a requirement of 95% D+1 deliveries. This approach works particularly well when the same delivery methods vary in quality depending on location and current load. The result is more predictable order fulfilment without manually “keeping an eye” on every parcel.

WooCommerce settings in the demo store: the “Poland” shipping zone with three delivery methods
Example Shipping zone in WooCommerce settings (demo store): courier, collection point pickup and free delivery from 200 zł

Controlling delivery costs with AI also involves eliminating the causes of problems before they turn into complaints and returns. The model can limit the share of carriers with a high rate of complaints or damage, even if at first glance they are cheaper, because in overall terms they increase servicing costs. The key is to treat carrier selection as a multi-criteria decision: cost should go hand in hand with on-time delivery and the level of complaint risk. As a result, optimisation is not reduced to the “lowest rate”, but to consistently delivering on the delivery promise.

FAQ

Frequently asked questions

What data is needed for effective personalisation in e-commerce?

At a minimum, you need behavioural data, the product catalogue, stock levels, price history, and customer data in the form of identifiers and segments. Consistent events such as view, add_to_cart and purchase, as well as the correct product_id, are also key.

Why do e-commerce recommendations often fail despite a good AI model?

Most often, the problem lies not in the model itself, but in errors on the events side, user identification and product catalogue quality. Duplicate event sends, a missing purchase or messy IDs can throw recommendations off.

How does AI improve the shopping experience on a store website?

AI matches content, product order and banners to user behaviour and context in real time. It can also personalise the PLP, hero banner or communication so that it is more relevant and less random.

Does AI personalisation in e-commerce need to take privacy and consent into account?

Yes, because personalisation without respecting consent and access controls increases the risk of losing trust and compliance issues. The article highlights, among other things, pseudonymisation, separating PII from behavioural data and integration with a CMP.

What tests show whether AI recommendations really work?

Deployment should be decided by online tests such as A/B or bandit tests, based on business KPIs. An offline result alone is not enough, which is why the impact on RPV, CVR, AOV and margin is also checked.

When does marketing automation with AI deliver the best results?

It delivers the best results when AI adapts not only the content, but also the timing, frequency and segment to customer behaviour. The article also underlines the value of scenarios such as win-back, CLV and churn prediction, and frequency capping control.

Contents