Indian Flag
Government Of India
A-
A
A+
ORGANISATION
ScholarGraph-Synthetic

ScholarGraph-Synthetic

Synthetic citation-network graph (1,200 papers, ~2,600 edges, 8 topic classes) for semi-supervised GNN node classification; companion GCN reaches ~84% test accuracy vs ~35% for a matched graph-blind MLP baseline, demonstrating genuine structural signal.

About Dataset

ScholarGraph-Synthetic is a fully synthetic citation-network dataset for graph-based semi-supervised node classification. Nodes represent papers, each assigned to one of eight research-topic classes. Edges represent citations, generated with a homophilous Stochastic Block Model: papers preferentially cite others in the same topic, producing the community structure characteristic of real citation graphs, with an average node degree comparable to the real Cora/Citeseer benchmarks. Node features are a 300-dimensional class-conditional bag-of-words vector, deliberately blended with a shared background vocabulary so that features alone are a weak classifier. This dataset ships a matched-pair ablation (identical features, split, and layer sizes, only the propagation mechanism differs) confirming that a graph-blind MLP baseline scores far lower than GCN/GAT, i.e. the graph structure is doing real work rather than being a redundant add-on. The dataset follows the standard low-label-rate transductive split convention: roughly 20 labeled training nodes per class, separate validation and test sets, and remaining nodes unlabeled but part of the graph. Contains no real bibliographic, personal, or proprietary data.

Purpose of Dataset

This Dataset Supports Research On Graph Neural Networks For Node Classification, Specifically Enabling A Rigorous Test Of Whether Graph Structure Genuinely Contributes Beyond Node Features Alone, Rather Than Assuming It Does. It Targets Researchers Comparing Gcn, Gat, And Other Graph-based Architectures Against Non-graph Baselines, And Demonstrates A Matched-pair Ablation Methodology (Identical Data And Split, Only The Model Architecture Varies) That Can Be Reused To Validate Graph-model Design Choices On Other Datasets. The Tunable Feature-separability Design Also Makes It Useful As A Controlled Testbed For Studying When Graph Structure Matters Versus When It Doesn't.

Activity Overview Activity Overview

  • Downloads0
  • Downloads 4
  • File Size 1.86 MB
  • Views 6

Tags Tags

  • graph

License Control License Control

Attribution 3.0 Unported (CC BY 3.0)

dataset.json ( 1.86 MB )


To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.

Data Quality Score BetaData Quality Score Beta

Version Control Version Control

FolderVersion 1(1.86 MB)
  • JAI MALI·9 day(s) ago
    • application/json
      dataset.json