TGraphX › Insights › Semi-Supervised Node Classification: TGraphX vs Traditional Approaches
← Back to Insights

Semi-Supervised Node Classification: TGraphX vs Traditional Approaches

Target keyword: semi-supervised node classification graph pytorch

Semi-Supervised Node Classification: TGraphX vs Traditional Approaches

Semi-supervised learning on graphs is one of the clearest examples of how graph structure can amplify small amounts of labeled data. A node classification task might have thousands of nodes but only a handful of labels per class. Traditional machine learning requires either accepting the small labeled set or investing in expensive labeling. Graph neural networks exploit the relational structure: labeled nodes influence their neighbors through message passing, and those neighbors influence their neighbors in turn, propagating label information across the graph.

This article examines semi-supervised node classification from multiple angles: classical label propagation, GCN as a semi-supervised method, self-training, and TGraphX implementations of each. It also discusses how to set up train/val/test masks and when each approach is appropriate.


What Semi-Supervised Learning on Graphs Means

In the standard semi-supervised graph setup, you have a single large graph with N nodes. A small fraction of nodes have labels (the training set). The rest are unlabeled at training time. The goal is to predict labels for all remaining nodes, including both the validation and test sets.

The key difference from a standard supervised classification problem is that the unlabeled nodes are not held out from the model — they are visible during training. The model sees the full graph topology and all node features, but only receives supervision signals from the labeled subset. Graph structure acts as a regularizer, encoding the assumption that connected nodes tend to share labels (homophily).


Prerequisites

This article assumes familiarity with:
- GNN message passing basics
- The TGraphX Graph object
- PyTorch training loops

For background on graph construction and feature handling, see the MNIST as a graph tutorial.


Setting Up Train/Val/Test Masks

A key design choice in semi-supervised graph learning is the masking setup. Rather than splitting nodes into separate datasets, you keep all nodes in one graph and use boolean masks to select which nodes contribute to the loss:

python
import torch
        from tgraphx import Graph
        
        # Example: Cora-style setup
        # 2708 nodes, 7 classes, 140 labeled training nodes (20 per class)
        num_nodes = 2708
        
        # Create masks
        train_mask = torch.zeros(num_nodes, dtype=torch.bool)
        val_mask = torch.zeros(num_nodes, dtype=torch.bool)
        test_mask = torch.zeros(num_nodes, dtype=torch.bool)
        
        # Assign 20 nodes per class to training (140 total for 7 classes)
        train_indices = torch.randperm(num_nodes)[:140]
        val_indices = torch.randperm(num_nodes)[140:640]   # 500 validation nodes
        test_indices = torch.randperm(num_nodes)[640:1640]  # 1000 test nodes
        
        train_mask[train_indices] = True
        val_mask[val_indices] = True
        test_mask[test_indices] = True
        
        g = Graph(
            node_features=torch.randn(num_nodes, 1433),
            edge_index=torch.load("cora_edge_index.pt"),
            node_labels=torch.randint(0, 7, (num_nodes,)),
        )
        

The training loop uses only the train_mask to compute loss, but all nodes participate in the forward pass (message passing sees the full graph):

python
logits = model(g.node_features, g.edge_index)  # [N, num_classes], all nodes
        loss = criterion(logits[train_mask], g.node_labels[train_mask])
        val_acc = (logits[val_mask].argmax(1) == g.node_labels[val_mask]).float().mean()
        

Approach 1: Label Propagation

Label propagation is the classical semi-supervised baseline. It spreads label distributions from labeled to unlabeled nodes by iteratively averaging neighbor labels:

python
from tgraphx.mining.label_prop import label_propagation
        
        # Returns predicted class probabilities for all nodes
        pred_probs = label_propagation(
            g,
            train_mask=train_mask,
            node_labels=g.node_labels,
            num_classes=7,
            num_iterations=50,
            alpha=0.85,  # teleportation parameter: controls label vs structure weight
        )
        
        predictions = pred_probs.argmax(dim=1)
        test_acc = (predictions[test_mask] == g.node_labels[test_mask]).float().mean()
        print(f"Label Propagation test accuracy: {test_acc:.4f}")
        

Label propagation requires no training phase and no GPU. It runs in seconds on graphs with tens of thousands of nodes. However, it does not use node features at all — only graph topology and the seed labels.


Approach 2: GCN as Semi-Supervised Learning

Kipf and Welling (2017) framed GCNs explicitly as a semi-supervised learning method. The key insight is that GCN's normalized aggregation is equivalent to a localized spectral convolution, acting as a smooth label-aware regularizer. The graph structure is used during both the forward pass (feature aggregation) and implicitly during training (the loss is computed only on labeled nodes, but gradients flow through the aggregation to influence all node representations).

python
import torch.nn as nn
        import torch.nn.functional as F
        from tgraphx.layers.sage import TensorGraphSAGELayer
        
        class SemiSupervisedGNN(nn.Module):
            def __init__(self, in_dim, hidden_dim, num_classes, dropout=0.5):
                super().__init__()
                self.conv1 = TensorGraphSAGELayer(in_dim, hidden_dim)
                self.conv2 = TensorGraphSAGELayer(hidden_dim, num_classes)
                self.dropout = dropout
        
            def forward(self, x, edge_index):
                x = F.relu(self.conv1(x, edge_index))
                x = F.dropout(x, p=self.dropout, training=self.training)
                x = self.conv2(x, edge_index)
                return F.log_softmax(x, dim=1)
        
        model = SemiSupervisedGNN(in_dim=1433, hidden_dim=64, num_classes=7)
        optimizer = torch.optim.Adam(model.parameters(), lr=0.01, weight_decay=5e-4)
        
        for epoch in range(200):
            model.train()
            optimizer.zero_grad()
            log_probs = model(g.node_features, g.edge_index)
            loss = F.nll_loss(log_probs[train_mask], g.node_labels[train_mask])
            loss.backward()
            optimizer.step()
        
            if epoch % 20 == 0:
                model.eval()
                with torch.no_grad():
                    log_probs = model(g.node_features, g.edge_index)
                    val_acc = (log_probs[val_mask].argmax(1) == g.node_labels[val_mask]).float().mean()
                print(f"Epoch {epoch}: val_acc = {val_acc:.4f}")
        

Approach 3: Self-Training

Self-training is an iterative approach: train a model on the labeled set, predict labels for high-confidence unlabeled nodes, add those pseudo-labeled nodes to the training set, and retrain. This is sometimes called pseudo-labeling.

python
def self_training_step(model, g, train_mask, confidence_threshold=0.95):
            model.eval()
            with torch.no_grad():
                log_probs = model(g.node_features, g.edge_index)
                probs = log_probs.exp()
                max_probs, pseudo_labels = probs.max(dim=1)
        
            # Find unlabeled nodes with high-confidence predictions
            unlabeled = ~train_mask
            high_confidence = (max_probs > confidence_threshold) & unlabeled
        
            # Add these to the training mask
            new_train_mask = train_mask.clone()
            new_train_mask[high_confidence] = True
        
            print(f"Added {high_confidence.sum().item()} pseudo-labeled nodes")
            return new_train_mask, pseudo_labels
        
        # Training loop with self-training
        current_train_mask = train_mask.clone()
        for iteration in range(3):
            # Train model on current labeled set
            for epoch in range(100):
                model.train()
                optimizer.zero_grad()
                log_probs = model(g.node_features, g.edge_index)
                loss = F.nll_loss(log_probs[current_train_mask], g.node_labels[current_train_mask])
                loss.backward()
                optimizer.step()
        
            # Expand training set with pseudo-labels
            current_train_mask, pseudo_labels = self_training_step(
                model, g, current_train_mask, confidence_threshold=0.95
            )
        

Self-training can improve accuracy significantly when the initial model is already reasonably good and the confidence threshold is set conservatively. When the initial model is poor, pseudo-labels can be noisy and self-training may hurt performance.


Using GIN for Semi-Supervised Classification

TGraphX's GIN layer can also be used for semi-supervised node classification, though GIN was originally designed for graph classification:

python
from tgraphx.layers.gin import TensorGINLayer
        
        class GINNodeClassifier(nn.Module):
            def __init__(self, in_dim, hidden_dim, num_classes):
                super().__init__()
                self.gin1 = TensorGINLayer(in_dim, hidden_dim, train_eps=True, use_batchnorm=True)
                self.gin2 = TensorGINLayer(hidden_dim, num_classes, train_eps=True, use_batchnorm=False)
        
            def forward(self, x, edge_index):
                x = F.relu(self.gin1(x, edge_index))
                return F.log_softmax(self.gin2(x, edge_index), dim=1)
        

Method Comparison

Method Uses Features Uses Structure Training Required Handles New Nodes Typical Accuracy Range
Label Propagation No Yes No Partially Baseline
GCN / SAGE (semi-supervised) Yes Yes Yes No (transductive) Strong
GAT Yes Yes (attention) Yes No (transductive) Strong
GIN Yes Yes Yes No (transductive) Strong
Self-Training Yes Yes Yes (iterative) Partially Depends on base model

"Typical accuracy range" is deliberately vague because performance depends heavily on the dataset, label rate, and implementation details. No generic numbers are provided here — consult the dataset-specific literature for concrete expectations.


Inductive vs Transductive Settings

The standard semi-supervised GNN setup is transductive: the test nodes are visible during training (the model sees their features and position in the graph), but their labels are withheld. This is different from inductive learning, where test nodes are entirely unseen during training.

For inductive semi-supervised learning (e.g., predicting labels for nodes in a new, previously unseen graph), GraphSAGE-style sampling is more appropriate. See the neighbor sampling for large graphs article for the inductive variant.


Limitations and Honest Notes

Transductive GNNs cannot be deployed directly to new graphs. If your application involves new nodes being added to the graph after training, the model requires retraining or an inductive adaptation.

Label propagation assumes strong homophily. On heterophilic graphs (where connected nodes tend to have different labels), label propagation will actively hurt performance by propagating misleading label information.

Self-training is sensitive to the threshold and iteration count. Confidence calibration in GNNs is known to be poor: high softmax probability does not reliably indicate correct prediction. Self-training with poorly calibrated models can compound errors.

The 140-label Cora setup is artificial. Many real-world semi-supervised scenarios have significantly different class balance, label noise, and graph structure than Cora. Performance on Cora benchmarks does not transfer directly to production graphs.

For reproducible semi-supervised experiments in TGraphX, use tgraphx.reproducibility to seed all random number generators before constructing masks. See the GNN research reproducibility guide for the complete reproducibility workflow.


Reproducibility in Semi-Supervised Experiments

The random split in a semi-supervised experiment has a large effect on reported accuracy. With only 140 labeled nodes in a 2708-node graph, the specific nodes selected for training can change results by several percentage points. This is why the Cora benchmark has a canonical fixed split that the community uses for comparison.

When creating your own splits, always:
1. Fix the random seed before generating masks
2. Report which split strategy you used (random, stratified, etc.)
3. Run multiple seeds and report mean ± standard deviation

python
from tgraphx.reproducibility import ReproducibilityContext
        
        ctx = ReproducibilityContext(seed=42)
        ctx.seed_all()
        
        # Now all torch.randperm calls are deterministic
        perm = torch.randperm(num_nodes)
        train_mask[perm[:140]] = True
        

Failing to fix the mask generation seed is one of the most common reproducibility oversights in GNN papers.


Measuring Semi-Supervised Performance Honestly

Semi-supervised accuracy depends on the label rate in a highly nonlinear way. A model that achieves 85% accuracy with 20 labels per class may only achieve 70% with 5 labels per class. When reporting results, always state:
- Number of labeled nodes per class (not just total)
- Whether the validation set is used for early stopping
- Whether test set performance was observed before finalizing the model

The Cora paper convention uses a fixed validation set for early stopping. Tuning on the test set — even by manual inspection — inflates reported accuracy. The GNN research reproducibility guide covers this in detail.


What This Article Builds On

This article assumes familiarity with basic GNN message passing and the TGraphX Graph object. The mask-based semi-supervised setup is a foundational pattern in graph learning; understanding it is prerequisite to more advanced topics like inductive generalization, heterophilic graph learning, and graph transformers. For the inductive variant where test nodes are unseen during training, see the neighbor sampling guide.

The label propagation approach covered in the "Approach 1" section is the classical non-neural baseline. It is covered in more depth in the articles hub at /articles/.


Frequently Asked Questions

What label rate is "semi-supervised"? There is no exact threshold, but the typical range in the academic literature is 1–20 labeled examples per class. With more than ~50 labels per class, standard supervised methods become competitive.

Does semi-supervised GNN learning work on heterophilic graphs? Poorly. GNN-based semi-supervised learning relies on the homophily assumption — connected nodes tend to have similar labels. On heterophilic graphs (like some web-of-links datasets), message passing can actively hurt accuracy compared to methods that do not aggregate over neighbors.

Should I use dropout in the semi-supervised setting? Yes. Dropout is especially important in the semi-supervised setting because the small labeled set makes overfitting easy. Standard values (0.3–0.6 on the first layer) apply.

How do I get the Cora dataset in TGraphX format? The tgraphx.datasets module provides standard benchmark datasets including Cora. Check the TGraphX documentation at https://github.com/arashsajjadi/TGraphX for the current dataset loading API.