<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Combined Fault Detection and Discrimination Strategy for Resource-Sensitive Platforms</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Richard McWilliam</string-name>
          <email>r.p.mcwilliam@durham.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philipp Schiefer</string-name>
          <email>philipp.schiefer@durham.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alan Purvis</string-name>
          <email>alan.purvis@durham.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Engineering and Computing Sciences</institution>
          ,
          <addr-line>Science Laboratories</addr-line>
          ,
          <institution>Durham University</institution>
          ,
          <addr-line>Durham, DH1 3LE</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <abstract>
        <p>This paper presents a combined fault detection and discrimination strategy for CMOS logic incorporating active resource mitigation and monitoring. The approach is demonstrated for a NOR gate using a dual redundant gate design with selective mitigation and analogue or digital detection. The potential bene ts of the approach are discussed with respect to resource awareness and management within ne-grained logic.</p>
      </abstract>
      <kwd-group>
        <kwd>Self-repair</kwd>
        <kwd>fault detection</kwd>
        <kwd>redundancy</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Fault detection and mitigation within CMOS logic structures is a long-standing
challenge that is seeing new a emphasis for nanoscale and printable electronics.
The possibility for intrinsic resource awareness and management without
obfuscating management at higher design levels is an attractive proposition but
requires new gate and transistor level strategies. This paper presents ongoing
work into a combined ne-grained redundancy and active mitigation approach
with minimal resource overhead that enables selective fault detection, masking
and discrimination close to the point of fault manifestation.
On-line fault strategies have been discussed at length for future nanoscale
electronics where massive redundancy concepts become feasible [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However,
resourcesensitive platforms typically involve more conservative duplicate gate and/or
interconnect structures combined with majority signal generation in order to mask
faults and prevent their manifestation at critical outputs. Practical examples
involving triple and quad redundancy are illustrated in Fig. 1a-b. Combined logic
interleaving and quad-transistor structures have also been investigated [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. While
the use of regular cell structures is attractive, typical methods incur between 3{8
times resource overhead and do not achieve fault detection or discrimination. It
could be argued that fault detection triggers may be generated within
quadtransistor majority logic but determination of the speci c fault location and its
2
      </p>
      <p>A Combined Fault Detection and Discrimination Strategy
c
type becomes abstracted by the internal process of converting critical faults to
sub-critical faults.</p>
      <p>
        Fine-grained fault tolerant strategies are beginning to feature in future
nanoscale CMOS logic design with the principal aim of combating manufacturing
defects. This includes psueo-CMOS redundancy, a simple example of which is
illustrated in Fig. 1c. Another approach is reported in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] wherein defects present
in either N-type or P-type networks invokes switched active pull up or pull-down
loads. In this case, however, defect detection is not part of the repair method
and instead would be provided by additional built-in self-test (BIST) logic and
possibly external test equipment. Hard-fault mitigation approaches have been
proposed that are based on active switching matrices [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. However, self-detection
is once again not included as a part of the strategy.
      </p>
      <p>
        Field-programmable gate arrays (FPGA) provide exible platforms
featuring con gurable cellular architectures that support full or partial con guration.
Since their total resource utilisation rarely approaches 100%, there are
opportunities to provision redundant resources for fault mitigation. Even so, it is not yet
clear how spare resources may be reallocated to support online fault detection
and discrimination without resorting to external supervisory hardware/software
as typi ed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. While solutions based on custom programmable architectures
have been proposed that aim to address this limitationby enabling dynamic
resource allocation [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] , fault detection is still achieved through data error detection
and correction (EDC) hardware that is abstracted from the hardware fault.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Proposed Method</title>
      <p>
        The proposed strategy relies upon an alternative method referred to here as
Stuck-At Fault Resilient (SAFR) design, wherein xed dual redundancy is
combined with a fault triggering mechanism [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. An example logic NAND gate
implemented by the SAFR approach in comparison to the standard NAND gate
design is shown in Fig. 2, where dual redundancy is employed within the P- and
N-type networks. This is in contrast to quad redundant strategies (Fig. 2c).
A Combined Fault Detection and Discrimination Strategy
3
The dual redundancy strategy permits masking of any single stuck-o fault and
selective fault triggers for stuck-on faults depending upon the state of the inputs.
Of particular note is the fact that fault discrimination is not retained when higher
redundancy factors are used i.e., triple- and quad-transistors. Hence, a resource
trade-o between fault masking capacity and fault identi cation is present in
this approach.
2.2
      </p>
      <p>
        Discrimination and Mitigation
Selective fault masking allows for the detection of stuck-on faults considered to
be critical due to potential high current ow between VDD and GND.
Examination of the gate output response under fault condition, summarised in Table 1,
shows that at there is at least one input combination that generates current ow
between VDD and VSS for every single stuck-on fault. This may be exploited
to achieve discrimination of fault type by monitoring current imbalance in the
CMOS network or else periodic exercising of the gate inputs via digital test. The
P- and N-networks are combined with the switching network for a NOR gate
implementation are shown in Fig. 3, which includes weak active pull-up/down
loads typically used for defect repair [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], but which are used here for selective
online fault discrimination.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Resource Awareness and Management</title>
      <p>Resource considerations will be important for emerging printable and nanoscale
electronics due to their di ering densities and scope for building redundancy
structures based upon multi-gate and/or sub-gate nano-structures. Resource
management extending to the ne-grained levels should be explored for both
defect tolerance and hard-fault mitigation. Combining the above approach with
weak active pull-up/down loads creates an e cient active mitigation
mechanism that, when further combined with dual redundancy within the P- and</p>
      <p>A Combined Fault Detection and Discrimination Strategy</p>
      <p>A
B</p>
      <p>T1
T3</p>
      <p>VDD
T2</p>
      <p>T4
(a)</p>
      <p>GND</p>
      <p>A+B
T5</p>
      <p>T6</p>
      <p>T7</p>
      <p>T8
k
r
o
w
t
e
N
P
k
r
o
w
t
e
N
N</p>
      <p>S1
S2
(b)</p>
      <p>S3
VDD
S4</p>
      <p>VDD
N-networks, creates further resource awareness opportunities in the presence of
faults. Once a fault has been detected, partial isolation proceeds by switching
to pseudo-NMOS or PMOS mode wherein the nature of the fault may be
further characterised. For example, assuming a stuck-at high fault occurring within
the P-network (Transistors T1-T4 in Fig. 3a), the location of the fault is not
known a-priori. The circuit may rst be switched to pseudo-NMOS mode
(setting switches S1 and S3 in Fig. 3b) and, due to the complimentary nature of
the design, a second analogue/digital test will would reveal the same fault
behaviour summarised in Table 1. However, depending on the value of the weak
pull-down resistance of transistor T9, the digital test may pass without error
and the adapted circuit may continue to be used in a degraded state.
Alternatively, the circuit may be switched into pseudo-PMOS mode (switches S2 and S4)
whereupon the error no longer persists. Hence the state of the P- and N-networks
may be individually ascertained. The reverse situation of a fault occurring within
the N-network would proceed in identical fashion as described above. At all times
stuck-at low fault events are intrinsically masked.
5</p>
      <p>An further extension of resource awareness concerns continual resource
monitoring in the presence of intermittent faults. For the above case of the
pseudoPMOS con guration being activated in response to a stuck-high fault within
the P-network, a further option would be to periodically switch to the
pseudoNMOS con guration and check the P-network response to determine whether
the fault persists. This serves two functions: rst, intermittent faults may be
handled in a graceful manner and with speci c knowledge of their locality.
Second, disappearance of the fault allows for restoration of the full CMOS network
and non-degraded performance.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>Fault detection and discrimination remains a fundamental challenge in resource
management for integrated fault mitigation. The proposed dual redundancy
SAFR method achieves a combination of fault discrimination between
stuckhigh/stuck-low fault events and selective masking, thus reserving active
mitigation for stuck-high faults. Fine-grained resource mitigation proceeds by
combining redundancy with weak pull-up/down networks. Ongoing work is investigating
further logic gate con gurations and functional logic built from such gates.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgement</title>
      <p>This work was supported by the UK EPSRC Centre for Innovative
Manufacturing in Through-life Engineering Services (EP/I033246/1).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Von</surname>
          </string-name>
          <string-name>
            <surname>Neumann</surname>
          </string-name>
          , \
          <article-title>Probabilistic logics and the synthesis of reliable organisms from unreliable components," Automata studies</article-title>
          , vol.
          <volume>34</volume>
          , pp.
          <volume>43</volume>
          {
          <issue>98</issue>
          ,
          <year>1956</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>J.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <surname>E</surname>
          </string-name>
          . Leung, L. Liu, and
          <string-name>
            <given-names>F.</given-names>
            <surname>Lombardi</surname>
          </string-name>
          , \
          <article-title>A Fault-Tolerant Technique Using Quadded Logic and Quadded Transistors,"</article-title>
          <source>IEEE Transactions on Very Large Scale Integration (VLSI) Systems</source>
          , vol. PP, no.
          <issue>99</issue>
          , pp.
          <volume>1</volume>
          {
          <issue>1</issue>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>M.</given-names>
            <surname>Ashouei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Chatterjee</surname>
          </string-name>
          , \
          <article-title>Recon guring CMOS as Pseudo N/PMOS for Defect Tolerance in Nano-Scale CMOS,"</article-title>
          <source>in 21st International Conference on VLSI Design</source>
          ,
          <year>2008</year>
          .
          <source>VLSID</source>
          <year>2008</year>
          ,
          <year>2008</year>
          , pp.
          <volume>27</volume>
          {
          <fpage>32</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>R.</given-names>
            <surname>Kothe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Vierhaus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Coym</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Vermeiren</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Straube</surname>
          </string-name>
          , \
          <article-title>Embedded Self Repair by Transistor and Gate Level Recon guration," in Design and Diagnostics of Electronic Circuits and systems</article-title>
          ,
          <source>2006 IEEE</source>
          ,
          <year>2006</year>
          , pp.
          <volume>208</volume>
          {
          <fpage>213</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J.</given-names>
            <surname>Emmert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Stroud</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Abramovici</surname>
          </string-name>
          , \
          <article-title>Online Fault Tolerance for FPGA Logic Blocks,"</article-title>
          <source>IEEE Transactions on Very Large Scale Integration (VLSI) Systems</source>
          , vol.
          <volume>15</volume>
          , no.
          <issue>2</issue>
          , pp.
          <volume>216</volume>
          {
          <issue>226</issue>
          ,
          <string-name>
            <surname>Feb</surname>
          </string-name>
          .
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>P.</given-names>
            <surname>Bremner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Samie</surname>
          </string-name>
          , G. Drag y, A. G. Pipe, G. Tempesti,
          <string-name>
            <given-names>J.</given-names>
            <surname>Timmis</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Tyrrell</surname>
          </string-name>
          , \
          <article-title>SABRE: a bio-inspired fault-tolerant electronic architecture,"</article-title>
          <source>Bioinspir. Biomim.</source>
          , vol.
          <volume>8</volume>
          , no.
          <issue>1</issue>
          , p.
          <fpage>016003</fpage>
          ,
          <string-name>
            <surname>Mar</surname>
          </string-name>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>P.</given-names>
            <surname>Schiefer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>McWilliam</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Purvis</surname>
          </string-name>
          , \
          <article-title>Fault Tolerant Quadded Logic Cell Structure with Built-in Adaptive Time Redundancy,"</article-title>
          <source>Procedia CIRP</source>
          , vol.
          <volume>22</volume>
          , pp.
          <volume>127</volume>
          {
          <issue>131</issue>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>