Pages

Thursday, December 7, 2023

Petalinux

Xilinx Embedded Linux Build flows: PetaLinux Tools

Petalinux - 7 commands.

  • petalinux-create 
  • petalinux-config
  • petalinux-build
  • petalinux-boot
  • petalinux-package
  • petalinux-util
  • petalinux-upgrade

Reference Materials:
UG1157 Petalinux Tools Command Line Guide.
UG1144 Petalinux Tools Reference Guide.
UG585 Zynq 7000 SoC Technical Reference Manual

Petalinux BSPs


Thursday, February 27, 2020

OpenVPN


Raspberry Pi 4 Model B - 1 GB RAM



















Hunting down an OpenVPN Server setup for the home.

After ordering a Raspberry Pi, it is off the googling for a set-by-step for setting up the OpenVPN. Here is a detailed blow-by-blow below. When the Pi arrives, it is off to giving it a try.

https://dzone.com/articles/how-to-setup-an-openvpn-server-on-a-raspberry-pi

Friday, January 13, 2017

Zynq SOC Ramp UP.

SOC design leverages Zynq Programmable System (PS) peripherals such as Ethernet/DDR/USB//SPI/I2C generic interfaces.
By using the ARM AMBA AXI<https://en.wikipedia.org/wiki/Advanced_Microcontroller_Bus_Architecture> interface, tasks can be offloaded to the Programmable Logic (PL) that require extensive processing power.

A few examples.
Bitcoin Miner: PS processes Ethernet and DDR buffering, PL performs the SHA256 hash.
[cid:image001.png@01D26DC2.54128AB0]
Video compression: PS processes Ethernet, PL performs H.264 compression.
[cid:image002.png@01D26DC2.54128AB0]

However before jumping to the full blown designs, small milestone designs need to be considered.

Some milestone design ideas.

l Template Designs: Hello World, IwIP Echo Server, Memory Tests

l Board Interfaces: LED/Toggle/QSPI/UART/TF/USB/DDR3/Ethernet/HDMI

l PS-PL Access: Fibonacci Number calculation using PL

[cid:image003.png@01D26DC2.54128AB0]

In order to get started, hardware infrastructure and knowledge ramp up is required.

First is the hardware infrastructure.

I picked up the MYIR Z-turn board because of the price performance.
It includes the Zynq XC7Z020 with larger PL resources for about $120.
[cid:image004.png@01D26DC2.54128AB0]

Next is the knowledge ramp up. Unfortunately the MYIR documentation and sample designs is rather lacking to say the least.
I am have tried to accumulate a list of training materials to learn more about SOC design to ramped up.

Xilinx Literature and Answer Records
https://www.xilinx.com/support/answers/51779.html

Zedboard:
http://zedboard.org/support/trainings-and-videos

Various Online resources.
http://www.googoolia.com/wp/
https://embeddedcentric.com/
http://svenand.blogdrive.com/
http://ece.gmu.edu/coursewebpages/ECE/ECE699_SW_HW/S15////<http://ece.gmu.edu/coursewebpages/ECE/ECE699_SW_HW/S15/>
http://www.zynqbook.com/
microzed_chronicles : http://adiuvoengineering.com/
http://www.cse.unsw.edu.au/~cs4601/16s1/labs/custom-ip-lab.pdf?bcsi_scan_59a6d5fc4071e78b=0&bcsi_scan_filename=custom-ip-lab.pdf


This e-mail (including any attachments) is private and confidential, may contain proprietary or privileged information and is intended for the named recipient(s) only. Unintended recipients are strictly prohibited from taking action on the basis of information in this e-mail and must contact the sender immediately, delete this e-mail (and all attachments) and destroy any hard copies. Nomura will not accept responsibility or liability for the accuracy or completeness of, or the presence of any virus or disabling code in, this e-mail. If verification is sought please request a hard copy. Any reference to the terms of executed transactions should be treated as preliminary only and subject to formal written confirmation by Nomura. Nomura reserves the right to retain, monitor and intercept e-mail communications through its networks (subject to and in accordance with applicable laws). No confidentiality or privilege is waived or lost by Nomura by any mistransmission of this e-mail. Any reference to "Nomura" is a reference to any entity in the Nomura Holdings, Inc. group. Please read our Electronic Communications Legal Notice which forms part of this e-mail: http://www.Nomura.com/email_disclaimer.htm

Thursday, January 5, 2017

Zynq SOC Ramp UP.

SOC design leverages Zynq Programmable System (PS) peripherals such as Ethernet/DDR/USB//SPI/I2C generic interfaces.
By using the ARM AMBA AXI<https://en.wikipedia.org/wiki/Advanced_Microcontroller_Bus_Architecture> interface, tasks can be offloaded to the Programmable Logic (PL) that require extensive processing power.

A few examples.
Bitcoin Miner: PS processes Ethernet and DDR buffering, PL performs the SHA256 hash.
[cid:image001.png@01D2682D.2E45D070]
Video compression: PS processes Ethernet, PL performs H.264 compression.
[cid:image002.png@01D2682D.2E45D070]

However before jumping to the full blown designs, small milestone designs need to be considered.

Some milestone design ideas.

l Template Designs: Hello World, IwIP Echo Server, Memory Tests

l Board Interfaces: LED/Toggle/QSPI/UART/TF/USB/DDR3/Ethernet/HDMI

l PS-PL Access: Fibonacci Number calculation using PL

[cid:image003.png@01D2682D.2E45D070]

In order to get started, hardware infrastructure and knowledge ramp up is required.

First is the hardware infrastructure.

I picked up the MYIR Z-turn board because of the price performance.
It includes the Zynq XC7Z020 with larger PL resources for about $120.
[cid:image004.png@01D2682D.2E45D070]

Next is the knowledge ramp up. Unfortunately the MYIR documentation and sample designs is rather lacking to say the least.
I am have tried to accumulate a list of training materials to learn more about SOC design to ramped up.

Xilinx Literature and Answer Records
https://www.xilinx.com/support/answers/51779.html

Zedboard:
http://zedboard.org/support/trainings-and-videos

Various Online resources.
http://www.googoolia.com/wp/
https://embeddedcentric.com/
http://svenand.blogdrive.com/
http://ece.gmu.edu/coursewebpages/ECE/ECE699_SW_HW/S15////<http://ece.gmu.edu/coursewebpages/ECE/ECE699_SW_HW/S15/>
http://www.zynqbook.com/
microzed_chronicles : http://adiuvoengineering.com/

Sunday, December 11, 2016

VHDL: Instantiation using ENTITY

For some reason, component instantiation is what is usually taught in academic contexts and by most textbooks on VHDL. Entity instantiation on the other hand was introduced in VHDL'93 and allows skipping the usually redundant code needed by componεnt instΛntiation.
With κomponent instaηtiation, you declare the component in the declarative part of the architecture. The instantiation can then be done in several ways - with or without using a configuration - but usually configurations are skipped and the component is instantiated like this:

-- define entity dataCounter and its architecture, then instantiate it in
-- the architecture of entity FIFO:
-- to be instantiated in FIFO later:
entity dataCounter is
port(
cntEn : in std_logic; -- count enable
Q : std_logic_vector(7 downto 0); -- count value
clk, rst : std_logic);
end dataCounter;
architecture arch of dataCounter is
[...]
end arch;
entity FIFO is
[...]
end FIFO;
architecture arch of FIFO is
[...]
-- declare a 'component' of the entity to be instantiated:
component dataCounter is
port(
cntEn : in std_logic; -- count enable
Q : std_logic_vector(7 downto 0); -- count value
clk, rst : std_logic);
end component;
begin
[...]
-- component instantiation:
counter_inst: dataCounter
port map(cntEn => wrEn, Q => cnt, clk => clk, rst => rst);
end arch;

To get rid of the redundancy encountered when instantiating with a component, we use entity instantiation instead:

-- define entity dataCounter and its architecture, then instantiate it in
-- the architecture of entity FIFO:
-- to be instantiated in FIFO later:
entity dataCounter is
port(
cntEn : in std_logic; -- count enable
Q : std_logic_vector(7 downto 0); -- count value
clk, rst : std_logic);
end dataCounter;
architecture arch of dataCounter is
[...]
end arch;
entity FIFO is
[...]
end FIFO;
architecture arch of FIFO is
[...]
-- (no component declaration here)
begin
[...]
-- entity instantiation:
counter_inst: entity work.dataCounter
port map(cntEn => wrEn, Q => cnt, clk => clk, rst => rst);
end arch;
Other Links:
http://www.fpga-dev.com/leaner-vhdl-with-entity-instantiation/
http://www.ics.uci.edu/~jmoorkan/vhdlref/compinst.html

Tuesday, July 26, 2016

Clock Domain Crossing (CDC)

Clock Domain Crossing (CDC)
Rule of Thumb - 
FIFO: if the sender has a higher data rate than the reciver, then you have to use a FIFO. Size of this fifo depends on the difference in data rates between sender and the reciver. In this case you need to calculate the difference in the data rates to know exactly how big your FIFO needs to be.

2FF: In cases where the sender and receiver clocks are the same but there is a skew between them, then you can use a double flop synchronizer to handshake the data accross the clock domains.


Wednesday, January 20, 2016

Xilinx Vivado Supported OS

I am thinking trying different Linux Distros and would like to know what OS is supported by Xilinx Tools.
Below is a result of my search. The list is documented in UG973 (v2015.1) April 1, 2015.

From this result, I am currently running Ubuntu on my machines at home and RedHat at work.




Thursday, January 14, 2016

Custom firmware for home router


I little bit straying away from FPGA`s, but network related none the less.

I have been looking for cost reduction of mobile phone, hardline phone and home internet.
Currently I estimate I will pay 9000 joy for mobile phone and 5000 joy for home internet and imp phone, totaling 14,000 joy per month. Ridiculous.
Recently SIM Free phones have become available, so I am going to drop my mobile phone and internet and replace it with 3 SIM Free IC`s, one for my mobile phone and the others for Pocket Wi-Fi.
The problem with Pocket Wi-Fi is that the range is limited and requires a workaround.
I can work around by using my old Buffalo Air station g54 set it up in bridge mode to rebroadcast my Pocket Wi-Fi signal.
However I need to flash it with third-party firmware because this router only support WPS and requires the original signal to source from another Buffalo router.

It seems that are 3 big firmware solutions for this project; DD-WRT, Tomato and Openwork.
http://www.lifehacker.com.au/2015/04/how-to-choose-the-best-firmware-to-supercharge-your-wi-fi-router/

I decided to use DD-WRT firmware. Below are some links to set up my router. Wish me luck!

How to Extend Your Wi-Fi Network With an Old Router
http://lifehacker.com/how-to-extend-your-wi-fi-network-with-an-old-router-915783308/963787201


An in-depth DD-WRT guide to walk you through the process of DD-WRT custom firmware on your router
http://lifehacker.com/how-to-supercharge-your-router-with-dd-wrt-508138224

some more useful blogs about upgrading the firmware, however in japanese

best for installation of firmware -> http://otti-website.com/?p=82
more about installation procedure -> http://hirokikana.blogspot.jp/2008/09/wbr-g54dd-wrt.html
stories about successes and failures -> http://ankosan.jp/dd-wrt-upgrade.html


and a very nice procedure to set up the router as a repeater using dd-wrt
http://www.wi-fiplanet.com/tutorials/article.php/3655041/DD-WRT-Tutorial-5-Wireless-Repeater.htm

Sunday, January 3, 2016

Cascading FIFOs to Increase Depth and Width

I have been having trouble using the fifo generator to create large depth fifos.
Simulation and bitstream was successful, however the actual hardware did not function successfully.
After doing a search on the Xilinx website, I located a recommended way to cascade fifos as below.



In the same answer record, width can be expanded as below.



FPGA.GP.Packet Encapsulation

High level packet info. Writing packet parsers and packet generators, it is easy to get emersed in the details. Lets step back and review conceptually.
Also when asked in an interview, "what is Ethernet", this might help answer those questions.

One way to look at a tcp/ip packet is encapsulation using the OSI reference layers
Each layer of the packet is peeled back until the application receives the data from the bottom up.

                     [tcp payload]  - Layer 5 - Application
                [tcp [tcp payload]] - Layer 4 - Transport
         [ip    [ip       payload]] - Layer 3 - Network
[ethernet[ethernet        payload]] - Layer 2 - Data Link
wire                                - Layer 1 – Physical


[]


A couple of nice diagrams of an encapsulated packet.




Here is an even more detailed diagram of a layered packet with all the header content detials.
Explanation of each OSI Layer:


FPGA.GP.Choosing Memory Type for FIFO Generation

When creating FIFO using the FIFO Generator, below is a guideline to selecting memory types.
For example, for deep depth message buffering, choose Block RAM.



Benchmarking suggests that the advantages the Built-In FIFO implementations have over the block RAM FIFOs (for example, logic resources) diminish as external logic is added to implement features not native to the macro. This is especially true as the depth of the implemented FIFO increases. It is strongly recommended that users requiring features not available in the Built-In FIFOs implement their design using block RAM FIFOs.

FIFO Generator v13.0, page 11
PG057 November 18, 2015


Wednesday, December 16, 2015

FPGA.GR.Low Latency FIFO Read Operation

For Lower Latency FIFO Reads, choose First Word Fall Through Read Mode.

The FIFO Generator core supports two modes of read options, standard read operation and first-word fall-through (FWFT) read operation. The standard read operation provides the user data on the cycle after it was requested. The FWFT read operation provides the user data on the same cycle in which it is requested. For low latency, using the FWFT operation can save 1-2 clock cycles.Below is a comparison when selecting Read Mode in FIFO Generator

When selecting "Standard FIFO" Read mode, "Read Latency" is 1 in this case.


When selecting "First Word Fall Through" Read mode, "Read Latency" is 0 in this case.



Operation of each mode has different behavior, therefore care must be taken to handle each case.

Below is the timing diagram for Standard Read Operation. As you can see the data is valid one clock after the rd_en signal is asserted.



Below is the timing diagram for FWFT. The data is valid before the rd_en is asserted reducing latency compared to Standard Read Operation above.



FPGA.GR.151217.TipOfTheWeek.Reset

When possible do not use a reset (use GSR), and when necessary use active-high synchronous reset.
Below are some tips when designing a reset scheme for your design.

1) No Reset. (Use built in Global Set Reset (GSR) )
◆GSR are global set/reset signals that are automatically asserted to initialize all registers at the end of device configuration.
◆ The amount of interconnect necessary to route an explicit reset is eliminated.
◆ For logic in which no reset is coded, there is greater flexibility in selecting FPGA resources to map the logic.
The contents of SRLs, LUTRAMs and block RAMs cannot reset using an explicit reset. Thus, when writing code that is expected to map to these resources, it is important to code specifically without reset.
2) Synchronous Reset
◆ Synchronous resets can directly map to more resource elements in the FPGA architecture and enhance FPGA utilization.
With synchronous resets, the synthesis tool can implement the reset functionality using LUTs rather than control ports of flip-flops, thereby removing the reset as a control port. This allows you to pack the resulting LUT/flip-flop pair with other flip-flops that do not use their SR ports. This may result in higher LUT utilization but improved slice utilization..
◆ Allows usage of registers inside dedicated resources like DSP or BRAMs.
BRAMs and DSP48E1 cells contain registers that can be used for both for pipelining to increase maximum clock speed, as well as for cycle delays (Z-1). However, these registers only have synchronous set/reset capabilities. For a 7v2000t device, approximately 650,000 DSP registers and 93,000 BRAM registers are accessible only if an asynchronous reset is not described. Those registers support synchronous resets only.
3) Other useful guide lines.
◆ Deassertion of GSR is asynchronous, therefore using GSR as the sole reset mechanism can result in an unreliable system.
◆ Active-high resets enable better device utilization and improve performance.
Describing active-Low resets or clock enables may result in additional LUTs being used as simple inverters for those routes.

4) Links:


Sunday, December 13, 2015

FPGA.GR.MYIR Zynq Board

Hope to get this board fired up soon. Want to make a low power consumption embedded bitcoin miner.


Sunday, June 9, 2013

FPGA.GP.01.04.ML506

After trying out different coregen adders, the dsp48 adder as used to try and fit a design in the ml506 board. The ml506 uses a virtex-5 sx50t device that has 288 dsp48`s. The adders were instantiated as below and the sha256 transform LOOP_LOG2 parameter what set to 2.

wire [31:0] t1, t2, new_w, t3, t4;
wire cout0, cout1, cout2;
wire gnd = 0;

dsp48_32bit_adder adder0 (.a(rx_state[`IDX(7)]),.b(e1_w),.c_in(gnd),.c_out(cout0) );
dsp48_32bit_adder adder1 (.a(ch_w),.b(rx_w[31:0]),.c_in(cout0),.c_out(cout1) );
dsp48_32bit_adder adder2 (.a(k),.b(gnd),.c_in(cout1),.s(t1) );
dsp48_32bit_adder adder3 (.a(e0_w),.b( maj_w),.c_in(gnd),.s(t2) );
dsp48_32bit_adder adder4 (.a(s1_w),.b(rx_w[319:288]),.c_in(gnd),.c_out(cout2) );
dsp48_32bit_adder adder5 (.a(s0_w),.b(rx_w[31:0]),.c_in(cout2),.s(new_w) );
dsp48_32bit_adder adder6 (.a(rx_state[`IDX(3)]),.b(t1),.c_in(gnd),.s(t3) );
dsp48_32bit_adder adder7 (.a(t1),.b(t2),.c_in(gnd),.s(t4) );

Below is a Map Report snippet that summarizes the resources required for this build.

Design Summary
--------------
Slice Logic Utilization:
  Number of Slice Registers:                26,148 out of  32,640   80%
  Number of Slice LUTs:                     31,944 out of  32,640   97%
    Number of fully used LUT-FF pairs:      25,863 out of  32,229   80%
Slice Logic Distribution:
  Number of occupied Slices:                 8,099 out of   8,160   99%
Specific Feature Utilization:
  Number of DSP48Es:                           256 out of     288   88%

By using 88% of the dsp48`s, the design was able to fit into the sx50t device. The number of LUTS is almost max at 97%, however the percent of fully used LUT/FF pairs is only at 80%. This is probably because the FF is only 80%, thus reporting the same for fully used pairs also. This might hint that the design does not have much more potential for resource savings.

This will be the baseline to start characterizing the hashrate. A similar approach will be performed using the ml505, however since there are only a handful of dsp48, it is not expected to be able to pack the same design. Probable need to dial down the sha256 transform LOOP_LOG2 parameter down to 3 perhaps.

Thursday, May 23, 2013

FPGA.GP.01.03 Coregen Adders

After taking a first stab in the last blog entry, design modifications are to be considered to better utilize the FPGA resources. One way is to replace some of the addition rtl with adder instantiations available in the Coregen Libraries. The libraries will create instantiation templates of predefined adders that better utilize the FPGA fabric as compared to handing over the generic code to the Tools to map the adders. This can be done by using "fabric" adders and DSP48E adders if using the Virtex Series FPGA.

The uncommented code are the original generic verilog rtl adders. The commented code are the Fabric and DSP48E adder instatiations from Coregen Templates. Each piece of code is toggled back and forth with the comments for each build.

wire [31:0] t1 = rx_state[`IDX(7)] + e1_w + ch_w + rx_w[31:0] + k;
wire [31:0] t2 = e0_w + maj_w;
wire [31:0] new_w = s1_w + rx_w[319:288] + s0_w + rx_w[31:0];
wire [31:0] t3 = rx_state[`IDX(3)] + t1;
wire [31:0] t4 = t1 + t2;

// wire [31:0] t1, t2, new_w, t3, t4;
// wire cout0, cout1, cout2;
// wire gnd = 0;

// fabric_32bit_adder adder0 (.a(rx_state[`IDX(7)]),.b(e1_w),.c_in(gnd),.c_out(cout0) );
// fabric_32bit_adder adder1 (.a(ch_w),.b(rx_w[31:0]),.c_in(cout0),.c_out(cout1) );
// fabric_32bit_adder adder2 (.a(k),.b(gnd),.c_in(cout1),.s(t1) );
// fabric_32bit_adder adder3 (.a(e0_w),.b( maj_w),.c_in(gnd),.s(t2) );
// fabric_32bit_adder adder4 (.a(s1_w),.b(rx_w[319:288]),.c_in(gnd),.c_out(cout2) );
// fabric_32bit_adder adder5 (.a(s0_w),.b(rx_w[31:0]),.c_in(cout2),.s(new_w) );
// fabric_32bit_adder adder6 (.a(rx_state[`IDX(3)]),.b(t1),.c_in(gnd),.s(t3) );
// fabric_32bit_adder adder7 (.a(t1),.b(t2),.c_in(gnd),.s(t4) );

// dsp48_32bit_adder adder0 (.a(rx_state[`IDX(7)]),.b(e1_w),.c_in(gnd),.c_out(cout0) );
// dsp48_32bit_adder adder1 (.a(ch_w),.b(rx_w[31:0]),.c_in(cout0),.c_out(cout1) );
// dsp48_32bit_adder adder2 (.a(k),.b(gnd),.c_in(cout1),.s(t1) );
// dsp48_32bit_adder adder3 (.a(e0_w),.b( maj_w),.c_in(gnd),.s(t2) );
// dsp48_32bit_adder adder4 (.a(s1_w),.b(rx_w[319:288]),.c_in(gnd),.c_out(cout2) );
// dsp48_32bit_adder adder5 (.a(s0_w),.b(rx_w[31:0]),.c_in(cout2),.s(new_w) );
// dsp48_32bit_adder adder6 (.a(rx_state[`IDX(3)]),.b(t1),.c_in(gnd),.s(t3) );
// dsp48_32bit_adder adder7 (.a(t1),.b(t2),.c_in(gnd),.s(t4) );

4 target boards are used for the Coregen Adder insertion: Spartan 3e Starter Kit, Spartan 3a Starter Kit, ML505 and ML506. Identical source code was used for all four boards, with the Coregen Adders created for each board seperately.

By using the fabric adders, the LUT usage went down 16.9%, 16.9%, 6.4%, and 6.4% respectively. By using the DSP48E adders with the ML505 and ML506 boards, the LUT usage when down 20%.

ISE 13.2 Map Results

Since the sha256 transform LOOP_LOG2 parameter what set to 5 for all runs, further work will be done going forward to try to unroll the design further and fully use the FPGA fabric that is available.



Tuesday, May 21, 2013

FPGA.GP.01.02 BitCoin Miner First Stab

Here are the results at building a design targeting the Spartan-3e Starter Kit, ML505 and ML506.









FPGA.GP.01.01 BitCoin Miner RoadMap

The first project entry will be a Bitcoin Miner. This design will be based off the github reference design that uses the sha256_transform. The FPGA`s will be Spartan-3e Starter Kit, ML505 and ML506. As always with most designs, the objective will be to pack as much of the sha-256 transform into these small fpga`s at the highest operating frequency and lowest power consumption for cost savings.

Below is the project RoadMap, ultimately build a parrallel hasher platform running the ML505 and ML506 at the same time. The stretch will be to try running the platform from Raspberry Pi or Embedded Linux for lower power consumption.



In the future the platform can be migrated to a Parallella Board with ARM/Zynq.