How does MarktMentor estimate sales figures for product research on bol?
Lars HurkmansCo-founder12 July 2023Temps de lecture 9 minutes
MarktMentor estimates sales figures on bol with its own algorithm instead of the shopping cart method. That model combines data sources such as the Retailer, Open and Advertising API with machine learning and regression models. For 80% of products, the estimate is within a 25% margin of error of actual sales.
Why did MarktMentor stop using the shopping cart method on bol?
MarktMentor stopped using the shopping cart method because it has structural limitations in accuracy, isn't scalable for tracking millions of products, and because intensive use of it clashes with bol's policy. Instead, we use a system we built ourselves, based on algorithmic estimates.
MarktMentor originally started, like many other tracking tools, by estimating sales figures using the well-known shopping cart method. This method involves checking a product's stock level at a given moment, then measuring the stock again at a later point. By looking at the difference, you can estimate how many units of a product were sold between the two measurement moments.
However, the shopping cart method has a number of limitations that make it difficult to always produce accurate estimates:
- Sales outside bol: Many sellers have an integrated stock system from which they manage all their sales. This means the stock level you see on bol is the seller's total stock level for that product. If sellers also sell this product outside bol, it's impossible to tell whether a sale happened on bol or elsewhere.
- Stock above 500: This is a well-known problem for tools that use the shopping cart method on bol. On bol you can't see stock levels above 500. This means it isn't possible to measure what happens once a product has more than 500 units in stock. There are also certain product categories (such as games) where you can only have a limited number of products in your shopping cart at once. This makes it very difficult to capture sales for these products based on stock information alone.
- Returns: When products are returned and subsequently found good enough to sell again, they're added back to stock. This means the same product can pass through the register multiple times. This leads to an overestimate of the total number of times a product has sold, especially for product groups with high return rates.
- Changes made by the seller: It isn't possible to tell whether changes in the shopping cart are the result of adjustments the seller made to their stock, just by looking at those changes. A seller might, for example, remove products from stock themselves, move them, or add products to stock.
- Multiple stock mutations happening in quick succession: It's possible for stock to be replenished at exactly the moment sales are also taking place. In such a situation, the net stock level stays unchanged, even though sales actually happened.
Although some of these problems can be minimized by measuring more frequently, that brings other challenges in terms of scalability. We went deeper into the technical limits of the shopping cart method later, in why the shopping cart method no longer works on bol.
Why isn't the shopping cart method scalable?
The shopping cart method isn't scalable because for hundreds of thousands or millions of products you'd need to measure stock multiple times a day, while bol sets limits on retrieving stock data.
If you want to track hundreds of thousands or even millions of products daily, measuring stock multiple times a day becomes challenging. This has to do with the size of the data, the complexity of the systems and, above all, the limits bol places on retrieving stock data from their website. In effect, we end up in a kind of cat-and-mouse game when we try to measure product stock levels on bol at scale and frequently. This is because of the restrictions bol imposes on scraping their website.
Scraping is a process in which automated programs "read" web pages and extract valuable information, such as product stock levels. To protect their website against overload or misuse, bol has introduced measures to counter excessive or intensive scraping. If we want to track millions of products daily, that requires frequent scraping, which can draw extra attention from bol. This results in a complex game in which we try to stay under the radar while still collecting the information we need. This can partly be achieved through techniques such as IP rotation or rate limiting, but those solutions are far from perfect.
So although it's technically possible to measure the stock of millions of products multiple times a day, bol's restrictions and the resulting cat-and-mouse game make this very challenging in practice. This means we may need to reduce the frequency of our stock measurements to avoid being blocked by bol. That inevitably affects the scalability and frequency of our stock measurements. In short: the shopping cart method becomes impractical once you want to track the sales of millions of products. For these and other reasons, we decided to stop using the shopping cart method and switch to a system we built ourselves, based on algorithmic estimates.
MarktMentor also places great value on a good relationship with bol. Bol tolerates scraping of their website to a certain extent, but intensive use of the shopping cart method, especially at scale, isn't appreciated. It can negatively affect their website's performance and hinder the experience of other users. We respect that position and follow their policy and wishes on this. Because we stick to these rules, we've also been able to become a bol Gold Partner.
These challenges around scalability and respecting bol's wishes, combined with the limitations of the shopping cart method, led us to decide to switch to an algorithmic approach for estimating sales figures. This allows us to produce estimates at a scale we could never achieve with the shopping cart method. We've compared how that scale relates to standalone product trackers in MarktMentor vs product trackers.
How does MarktMentor's algorithm work?
MarktMentor's algorithm is built on large amounts of data from multiple sources and uses it to estimate the historical sales of products through machine learning and regression models.
To gather that data, we need data sources. The data sources we use include bol's Retailer API, Open API and Advertising API. We combine the information from these sources with web scraping to obtain additional information.
Once the data has been collected, we start processing it thoroughly. This first means we need to "clean" the data, removing any errors or inconsistencies. We then prepare the data so it can be used in our algorithms.
Our next step is analyzing the data using advanced techniques such as machine learning and regression models. For this, we use, among others, the following variables:
- Historical popularity positions for the Netherlands & Belgium
- Product category
- Product offer information (price, delivery time) for the Netherlands & Belgium
- Day of the week
Based on these and other data sources, an estimate of historical sales is then produced. Here we use machine learning techniques that fall under the category of "supervised learning": we train our model on a dataset with real sales data, aiming to get parameters that can estimate how changes in the variables translate into a sales estimate. In general, this means that as the model sees more data, the sales estimates improve. We therefore also test the accuracy of our model against existing data.
To determine the accuracy of our system and improve it based on that, we look both at the product population as a whole and at conditioned subpopulations (determined based on the average historical sales of products). We use the accuracy based on the entire product population to optimize the general parameters, while the accuracy of the conditioned subpopulations is used to improve accuracy in specific situations. For the conditioned subpopulations, we look at the different order sizes of products (based on realized average historical sales) and optimize the parameters for each subpopulation.
How accurate is MarktMentor's data?
For 80% of products, the estimate is within a 25% margin of error of actual sales. No estimation model is perfect, so it's important to understand what that accuracy means.
Although our system is a significant improvement over the shopping cart method, it remains an estimate. For 80% of our products, we have a maximum margin of error of 25%. This means that for 80% of products, our sales estimate lies within a range of 25% of actual sales. This is best illustrated with an example.
Imagine you run an e-commerce store selling different kinds of products, such as clothing, electronics and books. You need to regularly estimate how much of each product you'll sell in the future to manage your stock effectively.
Those estimates aren't always perfect. Sometimes you estimate you'll sell 100 pairs of shoes, but you end up selling 80. There's then a difference between your estimate and actual sales, also called a margin of error. In this case, the margin of error is 20 pairs of shoes, or expressed as a percentage, 20% (because 20 of the 100 estimated pairs of shoes weren't sold).
You do this for every product you sell, which eventually gives you a whole list of margins of error for all your products. When you sort those margins of error from low to high, it turns out that for 80% of your products the margin of error is 25% or less. So for that 80% of your products, your sales estimate was within a range of 25% of actual sales.
If, for example, you sell 100 different products on your website, your estimate for 80 of those products was within 25% of actual sales. For the remaining 20 products, the difference between your estimate and actual sales was larger than 25%. This shows that the estimation system is fairly accurate for the majority of the products you sell.
What's the best way to interpret MarktMentor's sales data?
Use MarktMentor's sales data as an indication of demand for a product, not as an exact sales figure. Look at the data over longer periods and compare multiple similar products for a more general picture.
The sales data we show in our product research features are estimates. The algorithms are designed and optimized to give a good impression of a product's order size. So you can use the data to get a good indication of how much demand there is for products on bol.
For 80% of our products, the margin of error is limited to 25%. That means: in most cases, we come fairly close to actual sales figures. Do note: it can happen that for a specific product at a given moment, we deviate more than 25% from reality.
To get the most valuable insights, we advise you to examine the data in different ways:
- Look at sales figures over time: in the Product Overview you can view weekly sales estimates. This provides detailed information that helps you better understand a product's performance.
- Look at multiple similar products: with our Product Radar you can include various products in your analysis. By studying more products of a similar nature, you get a more general picture of that product type's sales performance.
Keep in mind that we update sales data daily. The information in the Product Database is therefore constantly changing, and the data in the Chrome extension is also updated every day. So don't focus solely on monthly figures, which can change significantly within a few days. Also look at data over longer periods, such as the past three months or the past year. We go further into how to weigh historical data in your product choices in historical data and product choice on bol.
What's the core of the switch from the shopping cart method to an algorithm?
MarktMentor has switched from the traditional shopping cart method to an algorithmic approach with advanced techniques such as machine learning and regression models. There's a certain margin of error in the estimates, but the new system is significantly better than the previous method. We keep refining our methods to offer valuable and accurate insights for product research on bol. We've summarized what to pay attention to when choosing a product research tool in choosing a product research tool for bol.