BehaviorGPT-Commerce 1
Purchase and session sequences modeled as the language of consumption.
- Parameters
- 150M
BehaviorGPT turns actions into real-time predictions.
An LLM predicts the next token from a fixed vocabulary. BehaviorGPT predicts the next action, where each event can carry text, imagery, a time, a place or a price, and the candidates form an open catalog.
One 12.5B-parameter model, pretrained on 150 billion user actions, evaluated on public benchmarks whose products and users it has never seen.
No training required. Fine-tuning is optional.
import unbox
client = unbox.UnboxAI(api_key="your_prod_key")
# the customer's live transaction from the auth stream
def authorize(event):
# 1. Load the cardholder's transaction history
user = client.history.load(user_id=event.card_hash)
# 2. Score how likely this transaction is fraud
result = client.score(history=user, action=event)
# 3. Act on the Unbox intelligence score
if result.probability > 0.95:
return "DECLINED: BEHAVIORAL_ANOMALY"
return "APPROVED"
import unbox
client = unbox.UnboxAI(api_key="your_prod_key")
# project a customer's value from their behavior
def lifetime_value(user_id):
# 1. Load the customer's full behavior history
user = client.history.load(user_id=user_id)
# 2. Predict expected value over the next 12 months
result = client.predict(history=user, target="clv_12m")
# 3. Act on the projected lifetime value
if result.value > 5000:
return "SEGMENT: HIGH_VALUE"
return "SEGMENT: STANDARD"
import unbox
client = unbox.UnboxAI(api_key="your_prod_key")
# assess default risk for a credit application
def underwrite(application):
# 1. Load the applicant's behavior history
user = client.history.load(user_id=application.user_id)
# 2. Score the probability of default
result = client.score(history=user, action=application)
# 3. Act on the Unbox risk score
if result.probability > 0.30:
return "DECISION: DECLINE"
return "DECISION: APPROVE"
import unbox
client = unbox.UnboxAI(api_key="your_prod_key")
# rank the next-best items for a live shopper
def recommend(session):
# 1. Load the shopper's behavior history
user = client.history.load(user_id=session.user_id)
# 2. Rank the catalog by predicted intent
result = client.rank(history=user, candidates=session.catalog)
# 3. Return the top items to surface
return result.top(k=10)
import unbox
client = unbox.UnboxAI(api_key="your_prod_key")
# re-rank search results by behavioral intent
def search(query, session):
# 1. Load the user's behavior history
user = client.history.load(user_id=session.user_id)
# 2. Retrieve and rank results for the query
result = client.search(history=user, query=query)
# 3. Return the ranked results
return result.top(k=20)
Purchase and session sequences modeled as the language of consumption.
Employee action sequences, predicting workforce dynamics from behavior rather than surveys.
The second commerce generation, learning taste from what people do rather than from pixels.
One backbone unifying behavior across retail and payments, transferring zero-shot across domains.
BehaviorGPT is a foundation model for human behavior. It learns directly from sequences of actions such as purchases, searches, and clicks, with no hand-engineered features. One pretrained model transfers across catalogs, companies, and tasks. The current version, BehaviorGPT-v4, has 12.5B parameters and was pretrained on 150 billion user actions.
A Large Behavioral Model (LBM) is a foundation model trained on chronological sequences of human actions instead of text. It predicts the next action from the actions before it, and what it learns transfers to tasks like fraud detection and recommendations. Unbox AI introduced the term, and BehaviorGPT is its LBM.
A recommender scores items against a user profile or matches similar users. BehaviorGPT models the sequence itself, so earlier actions change how the next one is read. That context lets it identify intent instead of scoring each action in isolation.
An LLM reads text and generates its answer token by token. BehaviorGPT reads the event sequence directly and ranks a whole catalogue in one forward pass: 0.7 ms per query, 48 times faster than a prompted LLM, and ahead of LLM baselines on accuracy in the BehaviorGPT-v4 benchmarks. Use an LLM for language, and BehaviorGPT for predicting what people do next.
Yes. The same action can be routine or suspicious depending on what preceded it, and BehaviorGPT models that sequence directly. It scores risk in real time as a session unfolds, and the model that powers recommendations adapts to fraud and risk without retraining from scratch.
Zero-shot, BehaviorGPT-v4 beats every baseline trained on up to 2.7 million samples of the target data, and one checkpoint leads all 13 public datasets it was tested on across retail, engagement, and payments. In production A/B tests, BehaviorGPT models have increased sales by double-digit percentages against incumbent systems, including +24% on recommendations and +16% search conversion.
Chronological event streams: transactions, searches, sessions, and interactions. No labels or engineered features are required.
Yes. Click Get API key, enter your work email, and your API key arrives by email together with a link to the demo and the example notebooks. Enterprise deployments are available separately.
The developer API key is free during early access. Enterprise deployments are priced on request: book a call.
Immediately. The key is shown on screen as soon as you sign up and emailed to you at the same time, so keep that email. If you sign up again with the same address, we point you back to the original email.
Yes. Our papers and documentation are on the research page. The BehaviorGPT-v4 abstract is on this page, and the full paper is coming soon.
A purchase history is a sequence, so we trained on it the way a language model trains on text: next-event prediction over roughly 600M online actions and 15B offline grocery purchases. What comes out is a single representation of a shopper that search, recommendations, and assortment can all read from, with none of the manual tagging the incumbent tools depend on.
Nothing about the recipe is specific to shopping, so we pointed it at workforce actions instead: 43M events from 80,000 employees, with salted IDs, jittered timestamps, and per-person splits. The architecture and hyperparameters carried over from commerce unchanged, and pretraining on next-event prediction beat training the same model from scratch on attrition by 7%.
The second commerce generation: 0.5B parameters trained on 215B interactions, or 4.7T tokens, across major art and design platforms. Similarity here is defined by behavior rather than pixels: which motifs people view, search for, and buy together decides what counts as alike, and that behavioral notion of taste beat the specialist engines it was A/B tested against.
Details coming soon.
PaperLanguage modeling had two defining moments. First, a single pretrained model, fine-tuned per task, displaced task-specific systems. Then scale made fine-tuning optional for many tasks, letting the same model handle them from few or no examples. Behavioral modeling learns from sequences of human actions, such as purchases, card payments, and app sessions. Unlike language modeling, it remains largely organized around individual datasets, tasks, and domains.
We present BehaviorGPT-v4, a 12.5B-parameter Large Behavioral Model pretrained by next-event prediction on 150 billion user actions (3.6 trillion tokens) from retail, engagement, and payments. The model represents behavior as chronological sequences of discrete, multimodal events, extending next-token prediction from a fixed vocabulary to an open set of candidates while still ranking them in a single step. On two public datasets whose catalogs and users were held out from pretraining, BehaviorGPT-v4 without any target-specific training outperforms every sequential, generative, transferable, tabular, and LLM baseline we evaluated, including specialists trained on the target's full training split. No baseline reaches BehaviorGPT-v4's zero-shot accuracy at any target-data budget we tested. Fine-tuning improves accuracy further, and the model gains accuracy from added target data 17 times faster than the strongest baseline. The same pretrained model serves three industries and 13 datasets with no per-dataset changes to architecture, objective, or vocabulary. Without adaptation, this one fixed model leads the strongest baseline on 30 of 31 dataset–task pairs; fine-tuned, it leads on all 31.
Compared with our 0.5B model, the 12.5B model transfers better zero-shot and generally needs less target data to reach a given accuracy. Because candidate representations are cached, the model ranks a full catalog in one forward pass, averaging under a millisecond per query in our evaluation harness, 48 times faster than the prompted LLM baseline. Production A/B tests of the BehaviorGPT model family provide complementary evidence, including double-digit improvements in conversion and orders in several deployments. We also document the recipe and design decisions behind the BehaviorGPT family, supported by internal experiments and ablations. We release an SDK and API for zero-shot deployment on new catalogs, an example notebook that reproduces a zero-shot evaluation through the API, and live demos built on the SDK, and we invite others to build their own.
Full paper coming soon.
Live demos
Storefront. Search and recommendations ranked by BehaviorGPT from each shopper's behavior. Open the demo