arxiv:2407.20060

RelBench: A Benchmark for Deep Learning on Relational Databases

Published on Jul 29

· Submitted by

zechengz on Aug 5

Upvote

Authors:

Kexin Huang ,

Jiaqi Han ,

Alejandro Dobles ,

Yiwen Yuan ,

Zecheng Zhang ,

Abstract

We present RelBench, a public benchmark for solving predictive tasks over relational databases with graph neural networks. RelBench provides databases and tasks spanning diverse domains and scales, and is intended to be a foundational infrastructure for future research. We use RelBench to conduct the first comprehensive study of Relational Deep Learning (RDL) (Fey et al., 2024), which combines graph neural network predictive models with (deep) tabular models that extract initial entity-level representations from raw tables. End-to-end learned RDL models fully exploit the predictive signal encoded in primary-foreign key links, marking a significant shift away from the dominant paradigm of manual feature engineering combined with tabular models. To thoroughly evaluate RDL against this prior gold-standard, we conduct an in-depth user study where an experienced data scientist manually engineers features for each task. In this study, RDL learns better models whilst reducing human work needed by more than an order of magnitude. This demonstrates the power of deep learning for solving predictive tasks over relational databases, opening up many new research opportunities enabled by RelBench.

View arXiv page View PDF Add to collection

Community

zechengz

Paper author Paper submitter 9 days ago

🌐 Website: https://relbench.stanford.edu/
📄 Paper: https://arxiv.org/abs/2407.20060
💻GitHub: https://github.com/snap-stanford/relbench

librarian-bot

8 days ago

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

puhsu

8 days ago

Cool! Very interested in the field and where it is going, good to see new benchmarks popping up. I also like that you compare GNNs with human performance in the experiments. The fact that many datasets do not overlap with the 4dbinfer https://arxiv.org/abs/2404.18209 is also nice (more datasets, good)

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

Upvote

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2407.20060 in a model README.md to link it from this page.

Datasets citing this paper 2

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.