<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>CITI'</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Course contrastive recommendation algorithm based on hypergraph convolution</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yanlie Zheng</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xueying Li</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Qingxia Shen</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fiberhome Telecommunication Technologies Co.</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wuhan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>China</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Hypergraph, Contrastive Learning, Recommendation, Course Learning 1</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>2</volume>
      <fpage>12</fpage>
      <lpage>14</lpage>
      <abstract>
        <p>With its powerful modeling ability of real data, graph-based convolutional recommendation algorithms have become an important tool for providing personalized services to users. However, the existing graph-based recommendation algorithms mainly face the following problems: 1. Simple graph-based neural network models are difficult to this complex interaction relationship between user-items and their interaction high order.2. Personalized recommendation systems rely on users to generate their own data, which makes it difficult to obtain effective labeling information.3. Interaction noise. User-item interaction is saturated with noise interference. In order to overcome these problems and improve the performance of personalized recommendation, in this paper, we design a course comparison recommendation method based on enhanced hypergraph convolution (DHSL-Cu). Specifically, first, we design the introduction of a dual hypergraph convolutional network to capture the higher-order relationships between users and items and the potential interaction information under different interaction types. Second, we design a momentum-driven twin-based architecture to optimize the target network using a momentum-based parameter update strategy. Finally, a negative sample selection mechanism based on course learning is designed as our training strategy. Finally, through extensive experiments on real-world datasets and parameter analysis, we demonstrate that our proposed DHSL-Cu model can efficiently improve recommendation services.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In recent years, due to the powerful learning capability demonstrated by graph
convolutional networks in the field of non-Euclidean data, many researchers have
introduced them into the recommendation domain [
        <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
        ]. They consider users and items in
the recommendation domain as nodes of an item, and the network relationships formed
between nodes as the interactions between users and items, and learn the embedded
information of users and items embedded in the graph structure through convolution [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Yin et al. proposed to demonstrate the potential of the graph CNN approach in Pinterest's
recommendation task [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. GC-MC [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and NGCF [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] have constructed bipartite graphs of
user-item interactions from user-item interaction data in order to learn user preferences
by utilizing the user-item graph structure. Compared to the above simple graph models in
recommendation, HyperGCN [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] combines hypergraph with graph neural network,
introduces the concept of hypergraph convolution [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and propagates the information of
nodes with the same interactions through hyperedge aggregation, which is able to
efficiently capture the higher-order interaction information between nodes.
      </p>
      <p>
        Although the above work has made some progress, it is still subject to the following 2
limitations: 1) the personalized recommendation system relies on users to generate their
own data, which makes it difficult to obtain effective labeling information [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. 2) It is well
known that GCN-based recommendation models are susceptible to the noise of user-item
interactions [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>For this reason, in this paper, in order to alleviate the shortcomings in existing
methods, we propose a course comparison recommendation method based on enhanced
hypergraph convolution (DHSL-Cu). Specifically, firstly, in order to capture the complex
interactions between users and programs, we introduce a dual hypergraph convolutional
model to serve as an encoder to capture the higher-order relationships between user
programs and the potential interaction information under different interaction types.
Second, we design a momentum-driven twin architecture to optimize the target network
using a momentum-based parameter update strategy. Next, we introduce the idea of
course learning and design a negative sample selection mechanism based on course
learning as our training strategy. Through extensive experiments on real-world datasets
as well as parameter analysis, the effectiveness and superiority of the DHSL-Cu model are
demonstrated.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Model</title>
      <p>1. Momentum-driven twin network based module: in this module both online network
and target network use the gated dual hypergraph based convolutional model as the
network structure. Meanwhile, we use the momentum update mechanism to optimize
the target network's representation learning module during the parameter training
process, which aims to utilize the characteristics of the momentum update mechanism
to encode historical information in the target network, thus guiding the online network
to learn to explore richer and more effective feature representations. At the same time,
the asymmetry of the comparison network is enhanced by this parameter updating
method thus alleviating the possible training collapse problem.
2. Negative sampling training strategy based on course learning, meanwhile, in order to
solve the problem of too much randomness of negative sampling in contrast learning,
we introduce the method of course learning. In the negative sampling process first
through the Scoring function S ( ) , we sort the negative samples from easy to hard. In
addition, a Pacing function P( ) is used to control the negative samples into the
training process. Then, the mutual and consistency information between the two views
is maximized, thus enhancing the user-item representation learning. Finally the
useritem feature representation learned through the whole model is used as the final
useritem representation for subsequent prediction tasks.</p>
      <sec id="sec-2-1">
        <title>2.1. Formalization</title>
        <p>Let a multipartite bipartite graph
  U , I, E 
j
set U and item set I ) consisting of edge sets , where E denotes
the jth type of edge. For example, Amazon data user behavior log can be represented as, a
multiple bipartite graph containing two kinds of nodes (users, items) and two kinds of
edges (clicks, queries).</p>
        <p>Input: Firstly, the original data is used to perform the corresponding data division and
processing, according to the interaction behavior between users and items, for the
embedding initialization of user-item nodes, to obtain the user-item initial feature matrix
X  XU , X I  , where XU  NF , X I  MF denote the feature representations of the
user set and the item set, respectively, and F is the feature dimensionality, N , M are the
have two different types of node sets (user
E  E1  E2  E K 
number of users and the number of items. The user eigenvalues
X</p>
        <p>I are used as inputs to the model in terms of hypergraph sets.</p>
        <p>Output: By feeding the corresponding data and structures to our model, the model
learns the final user-item representation for the next prediction recommendation task.</p>
        <p>U and item eigenvalues</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Hypergraph Building Blocks</title>
        <p>
          We base our approach on two different enhancement methods to enhance the topological
and attribute information of the graphs in order to obtain a new view of the user-item
interaction. For both views we construct the user hypergraph set GU for the online view
and the user hypergraph set G U' for the target view based on the user set U and the user
hypergraph set GI for the original view and the project hypergraph set GI' for the target
view based on the project set I . Among them, by combining the edge modification policy
[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] is used to construct the online view, and we use the GD policy [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] for constructing
the target view.
        </p>
        <p>We transform both views resulting from the above graph enhancement operations into
two isomorphic hypergraph sets. It is shown below:
=</p>
        <p>GU  gU,base , gU,1..., gU,k ,GI  gI,base , gI,1..., gI,k  ,
(3)
where gU, j  U ,U, j and gI, j  I , I, j , where  U , j and  I, j denote the
hyperedges in the hypergraphs gU , j and gI, j , respectively. Note that all hypergraphs in
GU share the same set of user nodes U while all hypergraphs in GI share the same set of
project nodes I . For each project node i  I , a hyperedge  U , j is introduced in the
hypergraph gU , j connecting u | u U , (u,i)  E j  , i.e., all user nodes in the set of user
nodes U that are directly connected to the project node i through the interaction type
E j . Similarly, for a user node u U , a hyperedge  I, j is introduced in gI, j that connects
i | i  I , (u, i)  E j  i.e., all project nodes in the set of project nodes I that are directly
connected to the user node u through the interaction type E j . Note that two special
 k   k 
hypergraphs gU ,base GU and gI,base GI are defined as g U , j1 U , j  and g V ,  
  j1 I, j  ,
i.e., the hypergraph is the one consisting of all interaction types between user items.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Dual Hypergraph Convolution Module</title>
        <p>Dual Hypergraph Convolution Module mainly utilizes two chi-sub hypergraphs GU and
GI in online view  or target view  to obtain the potential features of users and items
behind different interaction types by aggregating and propagating the feature
representations of users and items through the hyperedge aggregation and propagation of
corresponding hypergraphs constructed on the basis of different interaction types in two
hypergraphs through convolution operation. The specific process is as follows:</p>
        <p>First, we introduce an association matrix H in the hypergraph to describe the
relationship between nodes and hyperedges, given the hypergraph gU , j  U ,U , j ,
where j base,1,..., k denotes the different interaction types between user-items, and
k is the number of interaction types between user-items, and defines the association
matrix of gU , j as:</p>
        <p> 1 u  e  e U , j
HU, j u, e  </p>
        <p> 0 otherwise
where HU , j </p>
        <p>U U , j , U, j denotes the set of hyperedges in the hypergraph gU , j ,
where j base,1,..., k , By the same definition, we define the cross association matrix
for the hypergraph gI , j . Let the diagonal matrices DU , j u,u   |U||U| and BU , j  |U , j||U ,j|
denote the node degree matrix and the hyperedge degree matrix, respectively. Where
DU , j u,u    eU, j HU , j u, e and BU , j e,e   uU HU , j u,e . We utilize the hypergraph
spectral convolution operator to learn the embedding of each hypergraph in the model,
and the hypergraph convolution operator is denoted as:</p>
        <p>X l1   (HWH T  X l Pl )</p>
        <p>Then, we use the normalized hypergraph convolution operator and define the
hypergraph convolution operator for gU , j as:</p>
        <p>X Ul,1j  (DU1,j HU , jWU BU1,j HUT, j  X Ul , j PUl , j )
where  denotes the nonlinear activation function, X Ul , j  |U|dFl denotes the user
features in layer l , WU  |I||I| is a unitary matrix, PUl , j 
FlFl1 is a learnable
transformation matrix, and Fl and Fl1 denote the embedding dimensions in layers l and
l  1 . For the dual isomorphic hypergraph sets GU and GI , we independently learn the
node features from each isomorphic hypergraph gU , j and gI, j . As a result, we obtain
XU ,base , XU ,1,..., XU ,k  and X I ,base , X I ,1,..., X I ,k  .
node representations for users and projects based on different interaction types:</p>
        <p>Finally based on each layer of hypergraph convolution operator, after t rounds of
iterations, we can get the embeddings of users and items under different types of
interaction data, defined as XUt  XUt ,base , XUt ,1,, XUt ,k  , X Ut , j  R u Ft ,
(4)
(5)
(6)
X It   XIt,base , X It,1,, X It,k  , X It, j  R iFt . where Ft is the final embedding dimension
and k is the number of edge types. Finally, the feature values learned from multiple
interaction types are concatenated together to obtain the final embedding result as follows:
XU  X Ut WU  bU , X I  X It WI  bI
(7)
where WU ,WI  k1*Ft Ft and bU , bI  Ft are trainable parameters. XU  U Ft and
X I  I Ft . This allows us to obtain the final embedding Z  XU , X I  of users and items in
the online graph.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Momentum-driven course comparison based module</title>
      </sec>
      <sec id="sec-2-5">
        <title>2.4.1. Momentum-driven twinning-based network structure</title>
        <p>In this module we construct a momentum-driven twin network-based architecture as our
backbone, consisting of two identical network structures - the online network and the
target network - which use the same encoder - the hypergraph convolutional module.
Through the momentum optimization-based twin network feature, while the online
network encodes the node features, the historical training information is retained in the
target network, and then the mutual information between the two view representations is
maximized in the course-based comparative learning module, which guides the online
network to learn to explore richer and more effective feature representations.</p>
        <p>The details are as follows: the target network uses the momentum update mechanism
in the way of updating the parameters. The purpose is to make the target graph retain part
of the historical information in the optimization process through the characteristics of the
momentum update mechanism, and guide the online graph to learn richer and more
effective feature representations in the comparison learning session. The iterative process
of its model parameters is as follows:</p>
        <p>At  a  At1  1 a   At
(8)
items in the target graph is obtained Z  XU , X I .</p>
        <p>where a is an adjustable parameter controlling the degree of temporary information
retention, and At and At denote the learnable parameters of the encoder at round t for
the online network and target network, respectively. Thus, the embedding of users and</p>
      </sec>
      <sec id="sec-2-6">
        <title>2.4.2. Negative Sampling Module Based on Course Learning</title>
        <p>Aiming at the problems arising from the negative sampling strategies of existing
comparison learning papers, we design a novel negative sampling method based on course
learning. The main idea is to sort the negative samples according to the difficulty during
the training period, and introduce different negative samples into the training at different
stages according to the difference in the difficulty of the negative samples.</p>
        <p>S  zc   || zzvv || zzcc ||</p>
        <p>S (zc )  zv  zc</p>
        <p>S  zc   sim  zv , zc   (zv  zc )T .1(zv  zc )
where  is the covariance matrix of multidimensional random variables.</p>
        <p>
          In line with the literature [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], after we obtain the fraction S  zc  of each individual
negative embedding zc in the memory bank, we use the pacing function to schedule how
to introduce negative samples into the training process. The pacing function P(t)
specifies the size of the memory bank to be used at each step t . The memory bank for t
consists of the P(t) lowest scoring samples. Negative sample batches are sampled
uniformly from this set. We denote the complete memory bank size by C and the total
number of training steps by T . The formula is as follows:
Similar to the literature [39], firstly, for the embedding result of any node in the online
graph, c negative embeddings {zc} with different difficulties can be found in the target
graph. meanwhile, a scoring function S   is defined to map the different negative
embeddings into the target graph. that maps different negative embeddings to a numerical
score S  zc  to measure this difficulty. Further, the scoring function is set to sim  zv , zc  to
fully measure the difficulty of negative samples. The formula is as follows:
lNCE   log
        </p>
        <p>exp  sim  zv , zv  / 
exp(sim(zv , zv ) / )  Cc1exp(sim(zv , zc ) / )</p>
        <p>P(t)  1 .1log  t  e10   K
  T </p>
        <p>P(t)   t/T   K
where  is a smoothing parameter used to control the speed of the training process.
  1/ 2,1, 2 denote the root, linear and quadratic pacing functions, respectively.</p>
      </sec>
      <sec id="sec-2-7">
        <title>2.4.3. Contrastive loss</title>
        <p>To maximize the consistency between the online view and the target view. We use noise to
estimate the contrast loss. Specifically, we define a "memory bank" Q , which contains
each positive pair zv , zv , i.e., the same node in both views, and C negative samples
embedded in {zc}Cc1 , the different nodes in both views, and then we use the similarity
measure function sim  ,  to compute the positive pair zv , zv and negative pair{zv , zc}
, based on which the loss function is as follows:
(9)
(10)
(11)
(12)
(13)
(14)
where  denotes the temperature parameter. To simplify the calculation, we use dot
product as the similarity measure function.</p>
      </sec>
      <sec id="sec-2-8">
        <title>2.5. Model Optimization</title>
        <p>We first use the edges that already exist in the interaction data as positive edges and
extract some non-existing edges as negative edges. Finally, the loss function is designed by
maximizing the positive edge probability and minimizing the negative edge probability.
The loss function is defined as:</p>
        <p> n 
lr  (u,i)E   log (ZuT Zi )  (1  )  j1 Euj P(u) log(1 (ZuT Zuj ))  Eij P(i) log(1 (ZiT Zij ))(15)
where  is the sigmoid activation function,  is the weight parameter in order to
balance the importance of positive and negative samples, P(u) defines the distribution of
u candidate nodes, and n is the number of negative samples, the existing edges in the
multipartite bipartite graph are used as positive samples, and for each positive sample of
edge (u, i) , n negative edges are randomly sampled from node u and node i .</p>
        <p>Ultimately, we unify the representation learning module for the main recommendation
task and the self-supervised comparison learning module for auxiliary enhancement into
an overall learning framework. Formally, the final learning objective is defined as:
l  lr  lNCE
(16)
where  is a variable factor that controls the self-supervised contrast learning task.
Finally we train our model using the Adam algorithm.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental Comparison and Analysis</title>
      <sec id="sec-3-1">
        <title>3.1. Datasets</title>
        <p>
          To evaluate the performance of our model, we conducted experiments on two real-world
datasets, Amazon [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and Alibaba[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. The relevant attribute statistics of the two
datasets are shown in Table 1.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Experimental Setup</title>
        <p>Interaction Interaction Density</p>
        <p>
          types
60658 2 0.279%
27036 3 0.108%
Consistent with the literature [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], we use AUROC, AUPRC, Precision and Recall as
evaluation metrics. Meanwhile, we randomly select 60% of the edges as the training set
and the rest of the edges as the test set. The whole process is repeated five times to obtain
different random samples for the training and test sets. The mean and standard deviation
values of the classification evaluation metrics are reported.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Experimental Design</title>
        <p>In order to comprehensively test the performance of the DHSL-Cu model proposed in this
paper, we test it from different perspectives as follows.
1. comparison with mainstream advanced algorithms: in order to assess the effectiveness
of the DHSL-Cu model, this paper compares it with eight mainstream advanced
algorithms.
2. ablation analysis: the contribution of each component is analyzed.
3. parameter sensitivity analysis: the influence degree of momentum-driven update
weight ratio, learning rate, and self-supervised comparison learning factor is analyzed.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Baseline Algorithms</title>
        <p>
          To evaluate the performance of the DHSL-Cu model, we compare it with the following
existing mainstream state-of-the-art recommendation algorithms.
1. GraphSAGE [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]: proposes a generalized induction architecture that uses local
neighbor sampling of nodes and aggregated features to generate embeddings of nodes.
2. GCN [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] uses spectral graph convolution operators to learn local graph network
structures and node features in order to achieve semi-supervised learning directly on
graph structure data.
3. GAT [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] uses masked self-attention to assign different weights to each node and its
neighboring nodes based on their features, eliminating the need to use a
preconstructed graph.
4. HGNN [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] designed the hyperedge convolution operation to deal with data correlation
during representation learning. By this method, the hyperedge convolution operation
can be effectively utilized to capture the implicit layer representation of higher order
data structures.
5. HyperGCN [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] is a new method of GCN training for semi-supervised learning of
hypergraphs based on hypergraph theory.
6. DualHGCN [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] a self-supervised Dual Hypergraph Convolutional Network (DualHGCN)
model that transforms a multilayer two-part graph network into two sets of its
hypergraph sets.
7. SGL [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] designed three types of data augmentation based on different perspectives to
complement/supervise the recommendation task with self-supervised signals on
useritem graphs.
8. HCCF [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] designed hypergraph structure learning module and cross view hypergraph
contrast coding model based on contrast learning to learn better user representations
by characterizing both local and global collaborative relationships in joint embedding
space.
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Analysis of experimental results</title>
      </sec>
      <sec id="sec-3-6">
        <title>3.5.1. Comparison with baseline algorithm</title>
        <p>The results of all algorithm experiments on both datasets are shown in Table 2. From the
experimental results of the four evaluation metrics AUROC, AUPRC, Precision and Recall,
we can conclude that:
1. GraphSAGE, GCN and GAT perform poorly on the two datasets, which may be due
to the fact that these embedding methods based on simple homogeneous graphs
have a weak interaction representation and do not deal well with non-planar
relationships between nodes compared to hypergraph convolution.
2. DualHGCN outperforms HGNN and HyperGCN, probably because HyperGCN and
HGNN are both based on the same type of interaction information, and when
confronted with complex user-item interactions under different types, it is difficult
to capture the potential higher-order information of the user and the item based on
different types of interactions. While DualHGCN effectively captures the interaction
information between different interaction types by designing the information
transfer mechanism between hypergraphs, the overall performance of DualHGCN
is weaker than that of HCCF, probably because HCCF is designed with a
hypergraph-based comparison learning model, which efficiently utilizes the
localto-global cross-view supervision information.</p>
        <p>In conclusion, DHSL-Cu outperforms the other benchmarks. This may be related to the
following two main reasons: 1) DHSL-Cu can effectively incorporate historical training
information by designing a momentum-driven target network structure, and the
parameter updating method of momentum update with gradient vanishing can effectively
alleviate the training collapse problem during the self-supervised learning process. 2)
During the training process, we use negative sampling based on course learning, which is
different from the general random sampling method, and it can effectively utilize
local-toglobal cross-view supervision information. The negative sampling method based on
course learning is different from the general random sampling method, which can
effectively introduce different training phases according to different sample
characteristics, and effectively improve the generalization ability and prediction accuracy
of the model.</p>
      </sec>
      <sec id="sec-3-7">
        <title>3.5.2. Ablation Analysis</title>
        <p>In order to verify the different effects on the algorithm brought by the gating-based dual
hypergraph convolution module, the momentum-driven twin network structure, and the
negative sampling strategy based on course learning, we conducted experiments and
comparative analysis. The experimental results are shown in Table. 3, where DHSL-Cu-Init
denotes the feature initialization module only, DHSL-Cu-H introduces the gating-based
dual hypergraph convolution module after the feature priming module, and DHSL-Cu-M
introduces the momentum-driven twin-network-based module and comparative learning
on top of DHSL-Cu-H. DHSL-Cu-Both, i.e., the introduction of negative sampling strategy
after the introduction of the complete DHSL-Cu model.</p>
        <p>As shown in Table. 3, DHSL-Cu-Init performs the worst. DHSL-Cu-H better and
accurately shows the effectiveness and efficiency of the gating-based dual hypergraph
mechanism than it. The improvement in the performance of DHSL-Cu-M and
DHSL-CuBoth illustrates the power of the self-supervised contrast learning framework. Among
them, the best performance of DHSL-Cu-Both shows that the introduction of negative
sampling strategy can effectively improve and enhance the learning performance of
contrast learning, and DHSL-Cu and its variants outperform the Alibaba dataset on the
Amazon dataset. The possible reason could be that the Amazon dataset is denser, which
helps to capture more effective data during various types of user-item interactions.</p>
      </sec>
      <sec id="sec-3-8">
        <title>3.5.3. Parameter sensitivity analysis</title>
        <p>This section focuses on the effect of the factor  setting for self-supervised learning on
the model performance.</p>
        <p>As shown in Tables. 4, 5, the model performance at contrast learning factor less than
0.5 is improves with increasing contrast learning factor and reaches the optimal point for
all four evaluation metrics on the beta  0.5 two datasets, after which the performance
decreases with increasing contrast learning factor. This may be a result of learning loss
pairs that are too high in the gradient conflict between the prediction and comparison
tasks during training.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>We propose a new self-supervised learning method DHSL-Cu. Specifically, we first
generate two hypergraph views based on a two-part graph network of users and items
after 2 different graph enhancement strategies. For the target network in the twin
network we use a momentum-based parameter update mechanism. The slow moving
target network is made to encode the online network history observations. A course
comparison framework based on a negative sample selection mechanism for course
learning is also designed. Finally, the effectiveness and superiority of the proposed
DHSLCu model is well demonstrated through extensive experiments on real-world datasets as
well as parameter analysis.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Xie</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            <given-names>Y</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Przystupa</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kochan O. A</surname>
          </string-name>
          <article-title>Knowledge Graph Embedding Based Service Recommendation Method for Service-Based System Development</article-title>
          [J].
          <source>Electronics</source>
          ,
          <year>2023</year>
          ,
          <volume>12</volume>
          (
          <issue>13</issue>
          ):
          <fpage>2935</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Xu</surname>
            <given-names>X</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Przystupa</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kochan</surname>
            <given-names>O</given-names>
          </string-name>
          .
          <source>Social Recommendation Algorithm Based on SelfSupervised Hypergraph Attention [J]. Electronics</source>
          ,
          <year>2023</year>
          ,
          <volume>12</volume>
          (
          <issue>4</issue>
          ):
          <fpage>906</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Saito</surname>
            <given-names>Y</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yaginuma</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nishino</surname>
            <given-names>Y</given-names>
          </string-name>
          , et al.
          <article-title>Unbiased recommender learning from missingnot-at-random implicit feedback[C]//</article-title>
          <source>Proceedings of the 13th International Conference on Web Search and Data Mining</source>
          .
          <year>2020</year>
          :
          <fpage>501</fpage>
          -
          <lpage>509</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Zhu</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cui</surname>
            <given-names>P</given-names>
          </string-name>
          , et al.
          <article-title>Robust graph convolutional networks against adversarial attacks[C]//Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery &amp; data mining</article-title>
          .
          <year>2019</year>
          :
          <fpage>1399</fpage>
          -
          <lpage>1407</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Wang</surname>
            <given-names>X</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            <given-names>X</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            <given-names>M</given-names>
          </string-name>
          , et al.
          <source>Neural graph collaborative filtering[C]//Proceedings of the 42nd international ACM SIGIR conference on Research and development in Information Retrieval</source>
          .
          <year>2019</year>
          :
          <fpage>165</fpage>
          -
          <lpage>174</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Yadati</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nimishakavi</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yadav</surname>
            <given-names>P</given-names>
          </string-name>
          , et al.
          <article-title>Hypergcn: A new method for training graph convolutional networks on hypergraphs[C]//</article-title>
          <source>Advances in neural information processing systems</source>
          ,
          <year>2019</year>
          ,
          <volume>32</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Jiang</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kochan</surname>
            <given-names>O</given-names>
          </string-name>
          .
          <article-title>Fault diagnosis of RV reducer based on denoising time-frequency attention neural network [J]</article-title>
          .
          <source>Expert Systems with Applications</source>
          ,
          <year>2024</year>
          ,
          <volume>238</volume>
          :
          <fpage>121762</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Yanqiao</given-names>
            <surname>Zhu</surname>
          </string-name>
          , Yichen Xu,
          <string-name>
            <given-names>Feng</given-names>
            <surname>Yu</surname>
          </string-name>
          , et al.,
          <source>Graph Contrastive Learning with Adaptive Augmentation[C]// Proceedings of the Web Conference</source>
          <year>2021</year>
          , pp.
          <fpage>2069</fpage>
          -
          <lpage>2080</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Feng</surname>
            <given-names>Y</given-names>
          </string-name>
          ,
          <string-name>
            <surname>You</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            <given-names>Z</given-names>
          </string-name>
          , et al.
          <source>Hypergraph neural networks[C]//Proceedings of the AAAI conference on artificial intelligence</source>
          ,
          <year>2019</year>
          ,
          <volume>33</volume>
          (
          <issue>01</issue>
          ):
          <fpage>3558</fpage>
          -
          <lpage>3565</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Bian</surname>
            <given-names>Shuqing</given-names>
          </string-name>
          , Zhao Xin Wayne,
          <string-name>
            <given-names>Zhou</given-names>
            <surname>Kun</surname>
          </string-name>
          , et al.,
          <article-title>Contrastive Curriculum Learning for Sequential User Behavior Modeling via Data Augmentation[C]//</article-title>
          <source>Proceedings of the 30th ACM International Conference on Information &amp; Knowledge Management</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>3737</fpage>
          -
          <lpage>3746</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Qin</surname>
            <given-names>Xiuyuan</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuan</surname>
            <given-names>Huanhuan</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            <given-names>Pengpeng</given-names>
          </string-name>
          , et al.,
          <source>Meta-optimized Contrastive Learning for Sequential Recommendation[C]// Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2023</year>
          , pp
          <fpage>89</fpage>
          -
          <lpage>98</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Xue</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rajan</surname>
            <given-names>V</given-names>
          </string-name>
          , et al.
          <article-title>Multiplex bipartite network embedding using dual hypergraph convolutional networks[C]//</article-title>
          <source>Proceedings of the Web Conference</source>
          <year>2021</year>
          .
          <year>2021</year>
          :
          <fpage>1649</fpage>
          -
          <lpage>1660</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Wang</surname>
            <given-names>Shoujin</given-names>
          </string-name>
          , Hu Liang,
          <string-name>
            <surname>Wang Yan</surname>
          </string-name>
          , et al.,
          <source>Graph Learning based Recommender Systems: A Review[C]// Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence</source>
          ,
          <year>2021</year>
          ,pp.
          <fpage>4644</fpage>
          -
          <lpage>4652</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>